Search USGSSearch

Geology topics

Brian J. Knaus

Publications and source records attributed to Brian J. Knaus.

3 recordsLinked to original sources

SSR_pipeline: a bioinformatic infrastructure for identifying microsatellites from paired-end Illumina high-throughput DNA sequencing data

SSR_pipeline is a flexible set of programs designed to efficiently identify simple sequence repeats (e.g., microsatellites) from paired-end high-throughput Illumina DNA sequencing data. The program suite contains 3 analysis modules along with a fourth control module that can automate analyses of large volumes of data. The modules are used to 1) identify the subset of paired-end sequences that pass Illumina quality standards, 2) align paired-end reads into a single composite DNA sequence, and 3) identify sequences that possess microsatellites (both simple and compound) conforming to user-specified parameters. The microsatellite search algorithm is extremely efficient, and we have used it to identify repeats with motifs from 2 to 25bp in length. Each of the 3 analysis modules can also be used independently to provide greater flexibility or to work with FASTQ or FASTA files generated from other sequencing platforms (Roche 454, Ion Torrent, etc.). We demonstrate use of the program with data from the brine fly Ephydra packardi (Diptera: Ephydridae) and provide empirical timing benchmarks to illustrate program performance on a common desktop computer environment. We further show that the Illumina platform is capable of identifying large numbers of microsatellites, even when using unenriched sample libraries and a very small percentage of the sequencing capacity from a single DNA sequencing run. All modules from SSR_pipeline are implemented in the Python programming language and can therefore be used from nearly any computer operating system (Linux, Macintosh, and Windows).

Journal of Heredity

SSR_pipeline--computer software for the identification of microsatellite sequences from paired-end Illumina high-throughput DNA sequence data

SSR_pipeline is a flexible set of programs designed to efficiently identify simple sequence repeats (SSRs; for example, microsatellites) from paired-end high-throughput Illumina DNA sequencing data. The program suite contains three analysis modules along with a fourth control module that can be used to automate analyses of large volumes of data. The modules are used to (1) identify the subset of paired-end sequences that pass quality standards, (2) align paired-end reads into a single composite DNA sequence, and (3) identify sequences that possess microsatellites conforming to user specified parameters. Each of the three separate analysis modules also can be used independently to provide greater flexibility or to work with FASTQ or FASTA files generated from other sequencing platforms (Roche 454, Ion Torrent, etc). All modules are implemented in the Python programming language and can therefore be used from nearly any computer operating system (Linux, Macintosh, Windows). The program suite relies on a compiled Python extension module to perform paired-end alignments. Instructions for compiling the extension from source code are provided in the documentation. Users who do not have Python installed on their computers or who do not have the ability to compile software also may choose to download packaged executable files. These files include all Python scripts, a copy of the compiled extension module, and a minimal installation of Python in a single binary executable. See program documentation for more information.

Data Series

Taxonomic considerations in listing subspecies under the U.S. Endangered Species Act

The U.S. Endangered Species Act (ESA) allows listing of subspecies and other groupings below the rank of species. This provides the U.S. Fish and Wildlife Service and the National Marine Fisheries Service with a means to target the most critical unit in need of conservation. Although roughly one-quarter of listed taxa are subspecies, these management agencies are hindered by uncertainties about taxonomic standards during listing or delisting activities. In a review of taxonomic publications and societies, we found few subspecies lists and none that stated standardized criteria for determining subspecific taxa. Lack of criteria is attributed to a centuries-old debate over species and subspecies concepts. Nevertheless, the critical need to resolve this debate for ESA listings led us to propose that minimal biological criteria to define disjunct subspecies (legally or taxonomically) should include the discreteness and significance criteria of distinct population segments (as defined under the ESA). Our subspecies criteria are in stark contrast to that proposed by supporters of the phylogenetic species concept and provide a clear distinction between species and subspecies. Efforts to eliminate or reduce ambiguity associated with subspecies-level classifications will assist with ESA listing decisions. Thus, we urge professional taxonomic societies to publish and periodically update peer-reviewed species and subspecies lists. This effort must be paralleled throughout the world for efficient taxonomic conservation to take place.

Conservation Biology