Search USGSSearch

SEARCH · Search USGS

Results for “Bioinformatics”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

SSR_pipeline: a bioinformatic infrastructure for identifying microsatellites from paired-end Illumina high-throughput DNA sequencing data

SSR_pipeline is a flexible set of programs designed to efficiently identify simple sequence repeats (e.g., microsatellites) from paired-end high-throughput Illumina DNA sequencing data. The program suite contains 3 analysis modules along with a fourth control module that can automate analyses of large volumes of data. The modules are used to 1) identify the subset of paired-end sequences that pass Illumina quality standards, 2) align paired-end reads into a single composite DNA sequence, and 3) identify sequences that possess microsatellites (both simple and compound) conforming to user-specified parameters. The microsatellite search algorithm is extremely efficient, and we have used it to identify repeats with motifs from 2 to 25bp in length. Each of the 3 analysis modules can also be used independently to provide greater flexibility or to work with FASTQ or FASTA files generated from other sequencing platforms (Roche 454, Ion Torrent, etc.). We demonstrate use of the program with data from the brine fly Ephydra packardi (Diptera: Ephydridae) and provide empirical timing benchmarks to illustrate program performance on a common desktop computer environment. We further show that the Illumina platform is capable of identifying large numbers of microsatellites, even when using unenriched sample libraries and a very small percentage of the sequencing capacity from a single DNA sequencing run. All modules from SSR_pipeline are implemented in the Python programming language and can therefore be used from nearly any computer operating system (Linux, Macintosh, and Windows).

Journal of Heredity

Incomplete bioinformatic filtering and inadequate age and growth analysis lead to an incorrect inference of harvested-induced changes

Understanding the evolutionary impacts of harvest on fish populations is important for informing fisheries management and conservation and has become a growing research topic over the last decade. However, the dynamics of fish populations are highly complex, and phenotypes can be influenced by many biotic and abiotic factors. Therefore, it is vital to collect robust data and explore multiple alternative hypotheses before concluding that fish populations are influenced by harvest. In their recently published manuscript, Bowles et al, Evolutionary Applications, 13(6):1128 conducted age/growth and genomic analysis of walleye ( Sander vitreus ) populations sampled 13–15 years (1–2.5 generations) apart and hypothesized that observed phenotypic and genomic changes in this time period were likely due to harvest. Specifically, Bowles et al. (2020) documented differential declines in size-at-age in three exploited walleye populations compared to a separate, but presumably less-exploited, reference population. Additionally, they documented population genetic differentiation in one population pair, homogenization in another, and outlier loci putatively under selection across time points. Based on their phenotypic and genetic results, they hypothesized that selective harvest had led to fisheries-induced evolution (referred to as nascent changes) in the exploited populations in as little as 1–2.5 generations. We re-analyzed their data and found that (a) sizes declined across both exploited and reference populations during the time period studied and (b) observed genomic differentiation in their study was the result of inadequate data filtering, including retaining individuals with high amounts of missing data and retaining potentially undersplit and oversplit loci that created false signals of differentiation between time points. This re-analysis did not provide evidence for phenotypic or genetic changes attributable to harvest in any of the study populations, contrasting the hypotheses presented by Bowles et al. (2020). Our comment highlights the potential pitfalls associated with conducting age/growth analyses with low sample sizes and inadequately filtering genomic datasets.

Quebec

Attack of the PCR clones: Rates of clonality have little effect on RAD-seq genotype calls

Interpretation of high-throughput sequence data requires an understanding of how decisions made during bioinformatic data processing can influence results. One source of bias that is often cited is PCR clones (or PCR duplicates). PCR clones are common in restriction site-associated sequencing (RAD-seq) data sets, which are increasingly being used for molecular ecology. To determine the influence PCR clones and the bioinformatic handling of clones have on genotyping, we evaluate four RAD-seq data sets. Data sets were compared before and after clones were removed to estimate the number of clones present in RAD-seq data, quantify how often the presence of clones in a data set causes genotype calls to change compared to when clones were removed, investigate the mechanisms that lead to genotype call changes and test whether clones bias heterozygosity estimates. Our RAD-seq data sets contained 30%–60% PCR clones, but 95% of RAD-tags had five or fewer clones. Relatively few genotypes changed once clones were removed (5%–10%), and the vast majority of these changes (98%) were associated with genotypes switching from a called to no-call state or vice versa. PCR clones had a larger influence on genotype calls in individuals with low read depth but appeared to influence genotype calls at all loci similarly. Removal of PCR clones reduced the number of called genotypes by 2% but had almost no influence on estimates of heterozygosity. As such, while steps should be taken to limit PCR clones during library preparation, PCR clones are likely not a substantial source of bias for most RAD-seq studies.

Molecular Ecology Resources

How "simple" methodological decisions affect interpretation of population structure based on reduced representation library DNA sequencing: A case study using the lake whitefish

Reduced representation (RRL) sequencing approaches (e.g., RADSeq, genotyping by sequencing) require decisions about how much to invest in genome coverage and sequencing depth, as well as choices of values for adjustable bioinformatics parameters. To empirically explore the importance of these “simple” methodological decisions, we generated two independent sequencing libraries for the same 142 individual lake whitefish (Coregonus clupeaformis) using a nextRAD RRL approach: (1) a larger number of loci at low sequencing depth based on a 9mer (library A); and (2) fewer loci at higher sequencing depth based on a 10mer (library B). The fish were selected from populations with different levels of expected genetic subdivision. Each library was analyzed using the STACKS pipeline followed by three types of population structure assessment (FST, DAPC and ADMIXTURE) with iterative increases in the stringency of sequencing depth and missing data requirements, as well as more specific a priori population maps. Library B was always able to resolve strong population differentiation in all three types of assessment regardless of the selected parameters, largely due to retention of more loci in analyses. In contrast, library A produced more variable results; increasing the minimum sequencing depth threshold (-m) resulted in a reduced number of retained loci, and therefore lost resolution at high -m values for FST and ADMIXTURE, but not DAPC. When detecting fine population differentiation, the population map influenced the number of loci and missing data, which generated artefacts in all downstream analyses tested. Similarly, when examining fine scale population subdivision, library B was robust to changing parameters but library A lost resolution depending on the parameter set. We used library B to examine actual subdivision in our study populations. All three types of analysis found complete subdivision among populations in Lake Huron, ON and Dore Lake, SK, Canada using 10,640 SNP loci. Weak population subdivision was detected in Lake Huron with fish from sites in the north-west, Search Bay, North Point and Hammond Bay,showing slight differentiation. Overall, we show that apparently simple decisions about library construction and bioinformatics parameters can have important impacts on the interpretation of population subdivision. Although potentially more costly on a per-locus basis, early investment in striking a balance between the number of loci and sequencing effort is well worth the reduced genomic coverage for population genetics studies. More conservative stringency settings on STACKS parameters lead to a final dataset that was more consistent and robust when examining both weak and strong population differentiation. Overall, we recommend that researchers approach “simple” methodological decisions with caution, especially when working on non-model species for the first time.

Michigan, Ontario, Saskatchewan

Composition and distribution of fish environmental DNA in an Adirondack watershed

Background Environmental DNA (eDNA) surveys are appealing options for monitoring aquatic biodiversity. While factors affecting eDNA persistence, capture and amplification have been heavily studied, watershed-scale surveys of fish communities and our confidence in such need further exploration. Methods We characterized fish eDNA compositions using rapid, low-volume filtering with replicate and control samples scaled for a single Illumina MiSeq flow cell, using the mitochondrial 12S ribosomal RNA locus for taxonomic profiling. Our goals were to determine: (1) spatiotemporal variation in eDNA abundance, (2) the filtrate needed to achieve strong sequencing libraries, (3) the taxonomic resolution of 12S ribosomal sequences in the study environment, (4) the portion of the expected fish community detectable by 12S sequencing, (5) biases in species recovery, (6) correlations between eDNA compositions and catch per unit effort (CPUE) and (7) the extent that eDNA profiles reflect major watershed features. Our bioinformatic approach included (1) estimation of sequencing error from unambiguous mappings and simulation of taxonomic assignment error under various mapping criteria; (2) binning of species based on inferred assignment error rather than by taxonomic rank; and (3) visualization of mismatch distributions to facilitate discovery of distinct haplotypes attributed to the same reference. Our approach was implemented within the St. Regis River, NY, USA, which supports tribal and recreational fisheries and has been a target of restoration activities. We used a large record of St. Regis-specific observations to validate our assignments. Results We found that 300 mL drawn through 25-mm cellulose nitrate filters yielded greater than 5 ng/µL DNA at most sites in summer, which was an approximate threshold for generating strong sequencing libraries in our hands. Using inferred sequence error rates, we binned 12S references for 110 species on a state checklist into 85 single-species bins and seven multispecies bins. Of 48 bins observed by capture survey in the St. Regis, we detected eDNA consistent with 40, with an additional four detections flagged as potential contaminants. Sixteen unobserved species detected by eDNA ranged from plausible to implausible based on distributional data, whereas six observed species had no 12S reference sequence. Summed log-ratio compositions of eDNA-detected taxa correlated with log(CPUE) (Pearson’s R = 0.655, P < 0.001). Shifts in eDNA composition of several taxa and a genotypic shift in channel catfish ( Ictalurus punctatus ) coincided with the Hogansburg Dam, NY, USA. In summary, a simple filtering apparatus operated by field crews without prior expertise gave useful summaries of eDNA composition with minimal evidence of field contamination. 12S sequencing achieved useful taxonomic resolution despite the short marker length, and data exploration with standard bioinformatic tools clarified taxonomic uncertainty and sources of error.

New York

Assessing contaminants of emerging concern in the Great Lakes Ecosystem: A decade of method development and practical application

Assessing the ecological risk of contaminants in the field typically involves consideration of a complex mixture of compounds which may or may not be detected via instrumental analyses. Further, there are insufficient data to predict the potential biological effects of many detected compounds, leading to their being characterized as contaminants of emerging concern (CECs). Over the past several years, advances in chemistry, toxicology, and bioinformatics have resulted in a variety of concepts and tools that can enhance the pragmatic assessment of the ecological risk of CECs. The present Focus article describes a 10+- year multiagency effort supported through the U.S. Great Lakes Restoration Initiative to assess the occurrence and implications of CECs in the North American Great Lakes. State-of-the-science methods and models were used to evaluate more than 700 sites in about approximately 200 tributaries across lakes Ontario, Erie, Huron, Michigan, and Superior, sometimes on multiple occasions. Studies featured measurement of up to 500 different target analytes in different environmental matrices, coupled with evaluation of biological effects in resident species, animals from in situ and laboratory exposures, and in vitro systems. Experimental taxa included birds, fish, and a variety of invertebrates, and measured endpoints ranged from molecular to apical responses. Data were integrated and evaluated using a diversity of curated knowledgebases and models with the goal of producing actionable insights for risk assessors and managers charged with evaluating and mitigating the effects of CECs in the Great Lakes. This overview is based on research and data captured in approximately about 90 peer-reviewed journal articles and reports, including approximately about 30 appearing in a virtual issue comprised of highlighted papers published in Environmental Toxicology and Chemistry or Integrated Environmental Assessment and Management . Environ Toxicol Chem 2023;42:2506–2518. © 2023 SETAC. This article has been contributed to by U.S. Government employees and their work is in the public domain in the USA.

Environmental Toxicology and Chemistry

PumaPlex100: An expanded tool for puma SNP genotyping with low-yield DNA

The original PumaPlex is a high-throughput assay developed to genotype 25 single nucleotide polymorphisms (SNPs) in pumas ( Puma concolor ). Here, we describe the development of PumaPlex100 – an expanded version of the original assay that now genotypes > 100 SNPs. We tested 142 candidate SNPs and developed a panel of 101 polymorphic loci, which are spread across four multiplexes and suitable for genotyping of non-invasive samples. This panel will provide researchers a set of standardized markers, that can be analyzed with minimal bioinformatic skills, for the assessment of population structure and genetic diversity. These SNPs will serve as an important resource for the continued genetic monitoring of this species, especially monitoring through non-invasive sampling.

Sonora

Effects of carbamazepine to visual function in early life stage fish

The frequent detection of pharmaceuticals and personal care products (PPCPs) in the environment raises concern for aquatic systems. Carbamazepine (CBZ), an antiepileptic drug, is among the most detected PPCP globally, with concentrations in surface water exceeding those that induce toxicity to aquatic organisms. Non-targeted transcriptomic profiling was conducted in zebrafish ( Danio rerio ) larvae exposed to 0, 1, 5, 10, or 50 μg/L CBZ from 2 h post fertilization (hpf) through hatching, and then sampled at 48, 72, or 144 hpf. Transcriptomic profiles were annotated and characterized with in silico bioinformatic software to assess top enriched pathways and identify targets of environmentally relevant concentrations of CBZ and anchor molecular effects to higher levels of biological organization. Based on this analysis, CBZ was predicted to impair visual perception and sensory system development. The number of eye saccades, determined with a visually mediated behavioral assay, optokinetic response, was significantly reduced in 144 hpf larvae exposed to concentrations as low as 1 μg/L CBZ. These results indicate that environmentally relevant concentrations of CBZ may target and impact processes involved in visual function in fish.

Environmental Research

Attenuation of monkeypox virus by deletion of genomic regions

Monkeypox virus (MPXV) is an emerging pathogen from Africa that causes disease similar to smallpox. Two clades with different geographic distributions and virulence have been described. Here, we utilized bioinformatic tools to identify genomic regions in MPXV containing multiple virulence genes and explored their roles in pathogenicity; two selected regions were then deleted singularly or in combination. In vitro and in vivo studies indicated that these regions play a significant role in MPXV replication, tissue spread, and mortality in mice. Interestingly, while deletion of either region led to decreased virulence in mice, one region had no effect on in vitro replication. Deletion of both regions simultaneously also reduced cell culture replication and significantly increased the attenuation in vivo over either single deletion. Attenuated MPXV with genomic deletions present a safe and efficacious tool in the study of MPX pathogenesis and in the identification of genetic factors associated with virulence.

Virology

Phylogenetic techniques in geomicrobiology

Molecular biological techniques have revolutionized the field of geomicrobiology by providing researchers with robust techniques for identifying microorganisms and characterizing microbial communities in a wide variety of environments. These techniques have freed researchers from the constraints of classical culture-based microbiology and allowed the discovery of previously unknown phylogenetic diversity of microorganisms. In this chapter, we discuss the theory, methods, and workflow for applying molecular techniques to identify and characterize microbial populations. Our chapter focuses on SSU rRNA gene-based approaches, guiding the reader from sample collection and gene amplification through bioinformatics and statistical analysis. The workflow presented has been successfully used to identify microbial populations and community dynamics in a wide variety of habitats to understand the interactions between microbes and their environment.

Book chapter

Perfluorohexanesulfonic acid (PFHxS) induces hepatotoxicity through the PPAR signaling pathway in larval zebrafish (Danio rerio)

In recent years, the industrial substitution of long-chain per- and polyfluoroalkyl substances (PFAS) with short-chain alternatives has become increasingly prevalent, resulting in the widespread environmental detection of perfluorohexanesulfonic acid (PFHxS), a short-chain PFAS. However, there remains limited information about the potential adverse effects of PFHxS at environmental concentrations to wildlife. Here, early life stage zebrafish ( Danio rerio ) were exposed to environmentally relevant concentrations of PFHxS to better characterize the adverse effects of PFHxS on aquatic organisms. Nontargeted, transcriptomic analysis revealed potential hepatotoxic effects in exposed larvae, including macrovesicular and microvesicular hepatic steatosis, as well as focal liver necrosis. Morphological, histological, biochemical, and targeted transcript expression profiles further confirmed significant alterations in hepatocellular lesion numbers, liver pathological structures, relative liver size, liver biochemical parameters, and liver function genes. To validate the PPAR-mediated toxicological mechanism identified as an enriched pathway through in silico bioinformatics analysis, we tested the coexposure to an antagonist and PPAR morpholino knockdown. This intervention alleviated PFHxS-induced hepatic effects, including reductions in the levels of aspartate aminotransferase, alanine aminotransferase, total cholesterol, and total triglycerides. Our results demonstrate that environmentally relevant concentrations of PFHxS can impair liver development and function in fish, which could have potential risks to aquatic organisms.

Environmental Science & Technology

Toxicogenomics in regulatory ecotoxicology

Recently, we have witnessed an explosion of different genomic approaches that, through a combination of advanced biological, instrumental, and bioinformatic techniques, can yield a previously unparalleled amount of data concerning the molecular and biochemical status of organisms. Fueled partially by large, well-publicized efforts such as the Human Genome Project, genomic research has become a rapidly growing topical area in multiple biological disciplines. Since 1999, when the term “toxicogenomics” was coined to describe the application of genomics to toxicology (1), a rapid increase in publications on the topic has occurred (Figure 1). The potential utility of toxicogenomics in toxicological research and regulatory activities has been the subject of scientific discussions and, as with any new technology, has evoked a wide range of opinion (2–6).

Environmental Science & Technology

Shotgun sequencing of airborne eDNA achieves rapid assessment of whole biomes, population genetics and genomic variation

Biodiversity and its associated genetic diversity are being lost at an unprecedented rate. Simultaneously, the distributions of flora, fauna, fungi, microbes and pathogens are rapidly changing. Novel technology can help to capture and record genetic diversity before it is lost and to measure population shifts and pathogen distributions. Here we report the rapid application of shotgun long-read environmental DNA (eDNA) analysis for non-invasive biodiversity, genetic diversity and pathogen assessments from air. We also compared air eDNA with water and soil eDNA. Coupling long-read sequencing with established cloud-based biodiversity pipelines enabled a 2-day turnaround from airborne sample collection to completed analysis by a single investigator. To determine the full utility of airborne eDNA, we also conducted a local bioinformatic analysis and deep short-read shotgun sequencing. From outdoor air eDNA alone, comprehensive genetic analysis was performed, including population genetics (phylogenetic placement) of a charismatic mammal (bobcat, Lynx rufus ) and a venomous spider (golden silk orb weaver, Trichonephila clavipes ), and haplotyping humans ( Homo sapiens ) from natural complex community settings, such as subtropical forests and temperate locations. The rich datasets also enabled deeper analysis of specific species and genomic regions of interest, including viral variant calling, human variant analysis and antimicrobial resistance gene surveillance from airborne DNA. Our results highlight the speed, versatility and specificity of pan-biodiversity monitoring via non-invasive eDNA sampling using current benchtop/portable and cloud-based approaches. Furthermore, they reveal the future feasibility of scaling down (equipment and temporally) these approaches for near real-time analysis. Together these approaches can enable rapid simultaneous detection of all life and its genetic diversity from air, water and sediment samples for unbiased non-targeted information-rich genomics-empowered (1) biodiversity monitoring, (2) population genetics, (3) pathogen and disease-vector genomic surveillance, (4) allergen and narcotic surveillance, (5) antimicrobial resistance surveillance and (6) bioprospecting.

Nature Ecology & Evolution

Orthoptera-specific target enrichment (OR-TE) probes resolve relationships over broad phylogenetic scales

Phylogenomic data are revolutionizing the field of insect phylogenetics. One of the most tenable and cost-effective methods of generating phylogenomic data is target enrichment, which has resulted in novel phylogenetic hypotheses and revealed new insights into insect evolution. Orthoptera is the most diverse insect order within polyneoptera and includes many evolutionarily and ecologically interesting species. Still, the order as a whole has lagged behind other major insect orders in terms of transitioning to phylogenomics. In this study, we developed an Orthoptera-specific target enrichment (OR-TE) probe set from 80 transcriptomes across Orthoptera. The probe set targets 1828 loci from genes exhibiting a wide range of evolutionary rates. The utility of this new probe set was validated by generating phylogenomic data from 36 orthopteran species that had not previously been subjected to phylogenomic studies. The OR-TE probe set captured an average of 1037 loci across the tested taxa, resolving relationships across broad phylogenetic scales. Our detailed documentation of the probe design and bioinformatics process is intended to facilitate the widespread adoption of this tool.

Scientific Reports

Discovery of a unique Ig heavy-chain (IgT) in rainbow trout: Implications for a distinctive B cell developmental pathway in teleost fish

During the analysis of Ig superfamily members within the available rainbow trout (Oncorhynchus mykiss) EST gene index, we identified a unique Ig heavy-chain (IgH) isotype. cDNAs encoding this isotype are composed of a typical IgH leader sequence and a VDJ rearranged segment followed by four Ig superfamily C-1 domains represented as either membrane-bound or secretory versions. Because teleost fish were previously thought to encode and express only two IgH isotypes (IgM and IgD) for their humoral immune repertoire, we isolated all three cDNA isotypes from a single homozygous trout (OSU-142) to confirm that all three are indeed independent isotypes. Bioinformatic and phylogenetic analysis indicates that this previously undescribed divergent isotype is restricted to bony fish, thus we have named this isotype "IgT" (??) for teleost fish. Genomic sequence analysis of an OSU-142 bacterial artificial chromosome (BAC) clone positive for all three IgH isotypes revealed that IgT utilizes the standard rainbow trout VH families, but surprisingly, the IgT isotype possesses its own exclusive set of DH and JH elements for the generation of diversity. The IgT D and J segments and ?? constant (C) region genes are located upstream of the D and J elements for IgM, representing a genomic IgH architecture that has not been observed in any other vertebrate class. All three isotypes are primarily expressed in the spleen and pronephros (bone marrow equivalent), and ontogenically, expression of IgT is present 4 d before hatching in developing embryos. ?? 2005 by The National Academy of Sciences of the USA.

Proceedings of the National Academy of Sciences of

Assessing macroinvertebrate biodiversity in freshwater ecosystems: Advances and challenges in dna-based approaches

Assessing the biodiversity of macroinvertebrate fauna in freshwater ecosystems is an essential component of both basic ecological inquiry and applied ecological assessments. Aspects of taxonomic diversity and composition in freshwater communities are widely used to quantify water quality and measure the efficacy of remediation and restoration efforts. The accuracy and precision of biodiversity assessments based on standard morphological identifications are often limited by taxonomic resolution and sample size. Morphologically based identifications are laborious and costly, significantly constraining the sample sizes that can be processed. We suggest that the development of an assay platform based on DNA signatures will increase the precision and ease of quantifying biodiversity in freshwater ecosystems. Advances in this area will be particularly relevant for benthic and planktonic invertebrates, which are often monitored by regulatory agencies. Adopting a genetic assessment platform will alleviate some of the current limitations to biodiversity assessment strategies. We discuss the benefits and challenges associated with DNA-based assessments and the methods that are currently available. As recent advances in microarray and next-generation sequencing technologies will facilitate a transition to DNA-based assessment approaches, future research efforts should focus on methods for data collection, assay platform development, establishing linkages between DNA signatures and well-resolved taxonomies, and bioinformatics. ?? 2010 by The University of Chicago Press.

The Quarterly Review of Biology

Got acetylene: A personal research retrospective

In research, sometimes sheer happenstance and serendipity make for an unexpected discovery. Once revealed and if interesting enough, such a finding and its follow-up investigations can lead to advances by others that leave its originators ‘scooped’ and mulling about what next to do with their unpublished data, specifically what journals could it still be published in and be perceived as original. This is what occurred with us nearly 40 years ago with regard to our follow-up observations of acetylene fermentation and led us to concoct a ‘cock-and-bull’ story. We hypothesized about a plausible role for acetylene metabolism in the primordial biogeochemistry of Earth and the possibility of acetylene serving as a key life-sustaining substrate for alien microbes dwelling in the orbs of the outer solar system. With the passage of time, advances were made in whole-genome sequencing coupled with major in silico progress in bioinformatics. In parallel came the results of explorations of the outer solar system (i.e. the Cassini mission to Saturn and its moons). It now appears that these somewhat harebrained ideas of ours, arisen at first out of a sense of desperation, actually ring true in fact, and particularly well in song: ‘Tell a tale of cock and bull , Of convincing detail full Tale tremendous, Heav'n defend us! What a tale of cock and bull!' From ‘The Yeoman of the Guard’ by Gilbert & Sullivan .

FEMS Microbes

Identification of a novel arsenite oxidase gene, arxA, in the haloalkaliphilic, arsenite-oxidizing bacterium alkalilimnicola ehrlichii strain MLHE-1

Although arsenic is highly toxic to most organisms, certain prokaryotes are known to grow on and respire toxic metalloids of arsenic (i.e., arsenate and arsenite). Two enzymes are known to be required for this arsenic-based metabolism: (i) the arsenate respiratory reductase (ArrA) and (ii) arsenite oxidase (AoxB). Both catalytic enzymes contain molybdopterin cofactors and form distinct phylogenetic clades (ArrA and AoxB) within the dimethyl sulfoxide (DMSO) reductase family of enzymes. Here we report on the genetic identification of a “new” type of arsenite oxidase that fills a phylogenetic gap between the ArrA and AoxB clades of arsenic metabolic enzymes. This “new” arsenite oxidase is referred to as ArxA and was identified in the genome sequence of the Mono Lake isolate Alkalilimnicola ehrlichii MLHE-1, a chemolithoautotroph that can couple arsenite oxidation to nitrate reduction. A genetic system was developed for MLHE-1 and used to show that arxA (gene locus ID mlg _ 0216 ) was required for chemoautotrophic arsenite oxidation. Transcription analysis also showed that mlg _ 0216 was only expressed under anaerobic conditions in the presence of arsenite. The mlg _ 0216 gene is referred to as arxA because of its greater homology to arrA relative to aoxB and previous reports that implicated Mlg_0216 (ArxA) of MLHE-1 in reversible arsenite oxidation and arsenate reduction in vitro . Our results and past observations support the position that ArxA is a distinct clade within the DMSO reductase family of proteins. These results raise further questions about the evolutionary relationships between arsenite oxidases (AoxB) and arsenate respiratory reductases (ArrA). Arsenic is toxic to most organisms and is known to cause cancer in humans. However, bacteria have adapted several biotransformation pathways that function to either couple the reduction or oxidation of arsenicals to energy conservation and growth (1). The enzymologies of these two pathways have several features in common. The arsenate respiratory reductase (ArrAB) and arsenite oxidase (AoxAB) enzymes are usually composed of at least two subunits, a small iron-sulfur cluster-containing subunit (ArrB and AoxA) and a larger molybdopterin-containing catalytic subunit (ArrA and AoxB). Although they catalyze arsenic redox chemistry, ArrA and AoxB form distinct phylogenetic clades within the dimethyl sulfoxide (DMSO) reductase family of molybdenum-containing enzymes (16, 24). Culture-dependent approaches have resulted in the isolation of a variety of diverse bacteria that metabolize arsenic (reviewed in reference 26). Many of these isolates have had their genomes sequenced, which has been insightful for understanding the composition and diversity of arr and aox gene clusters. In the arsenite-oxidizing nitrate reducer Alkalilimnicola ehrlichii strain MLHE-1 (a haloalkaliphile isolated from Mono Lake [CA]) (10, 15), bioinformatic analysis of its genome revealed the absence of genes homologous to the arsenite oxidase genes of the aoxB type. Instead, two genes ( mlg _ 0216 and mlg _ 2426 ) were identified that better resembled the catalytic subunit of the arsenate respiratory reductase (20); however, MLHE-1 has not been shown to respire (or reduce) arsenate (15). Recent work by Richey et al. (20) showed that the Mlg_0216 protein (and not Mlg_2426) was expressed under chemolithoautotrophic (10 mM arsenite and 10 mM nitrate) growth conditions. Moreover, it was shown that Mlg_0216 exhibits both arsenate reductase and arsenite oxidase activities in vitro . These observations raised the question, is the mlg _ 0216 gene required for arsenite oxidation in vivo ? In this report, we addressed this question by developing a genetic system in MLHE-1, generating strains with mutations in mlg _ 0216 and mlg _ 2426 , and physiologically characterizing the resulting strains. Our results implicate mlg _ 0216 in chemolithoautotrophic arsenite oxidation coupled to nitrate respiration.

Journal of Bacteriology