Sunday, May 12, 2013

Sneak Peek: Elucidating the Effects of the Deep Water Horizon Oil Spill on the Atlantic Oyster Using RNA-Sequencing Data Analysis Methods

Join us this Tuesday, May 21st at 10 AM Pacific Time / 1:00 PM Eastern Time, for an interesting webinar on the effects of the Deep Water Horizon oil spill.

Speakers:
Natalia G. Reyero, PhD. – Mississippi State University
N. Eric Olson, PhD. – PerkinElmer Sr Leader Product Development

The Deep Water Horizon oil spill exposed the commercially important Atlantic oyster to over 200 million gallons of spill-related contaminants. To study toxicity effects, we sequenced the RNA of oyster samples from before and after the spill. In this webinar, we will compare and contrast the different data analysis methodologies used to address the challenge of an organism lacking a well-annotated genome assembly. Furthermore, we will discuss how the newly generated information provided insight into underlying biological effects of oil and dispersants on Atlantic oysters during the Deep Water Horizon oil spill.

REGISTER HERE to attend.

Thursday, February 28, 2013

ABRF 2013

The annual Association of Biomedical Research Facilities begins this weekend (March 2 - 5).  We [PerkinElmer] will be busy at the conference as participants and as a vendor supporting this great organization and our many customers. From client presentations to our own work we will share our latest and greatest.

Highlights:

Saturday: 3/2 "Breaking the Data Analysis Bottleneck: Solutions That Work for RNA and Exome Sequencing." Rebecca Laborde, Mayo Clinic will present on how the teams she works with use GeneSifter for they're NGS data analysis. This is part of Satellite Workshop 1: Applications of NGS.

Monday: 3/4 "Oyster Transcriptome Analysis by Next Gen Sequencing." Natalia Reyero, Genetics and Development Biology Center, NHLBI will give a presentation based on her award nominated poster (below) during the (RG5) Genomics Research Group. GeneSifter had a role in the data analysis.

Saturday and Monday Posters:

#7 "Identifying Mutations in Transcriptionally Active Regions on Genomes Using Next Generation Sequencing." Eric Olson, PerkinElmer, presents ways in which RNA-seq can be used to define transcripts to identify functional mutations in organisms that have sparsely annotated reference genomes. 

#11 "What Does It Take to Identify the Signal from the Noise in Molecular Profiling of Tumors?" Eric Olson, PerkinElmer presents ways to use RNA sequencing and bioinformatic approaches to filter the vast numbers of variants observed in DNA sequence data obtained from tumors to a manageable number that are most likely to be the drivers of tumor growth.

Award Nominee
#119 "Elucidating the Effects of the Deepwater Horizon Oil Spill on the Atlantic Oyster Using Global Transcriptome Analysis" Natalia Reyero, Genetics and Development Biology Center, NHLBI. If you are interested in learning about the aftermath of the Gulf of Mexico oil spill, you will find Natalia's work interesting. 

And that's not all. The the booth will be hopping. We will have meet the speaker opportunities on Sunday and Tuesday and as well as many demos of our GeneSifter Analysis and LIMS products.  PerkinElmer Informatics will have individuals to show new features in PerkinElmer's Electronic Laboratory Notebook and other products, and Caliper and Chemagen reps will be on hand to talk about the great things we do for sample prep.  

Check us out a Booth 522 to get schedules and see what's new.

Saturday, February 9, 2013

Genomics Genealogy Evolves

ResearchBlogging.org
The ways massively parallel DNA sequencing can be used measure biological systems is only limited by imagination. In science, imagination is an abundant resource.

The November 2012 edition of Nature Biotechnology (NBT) focused on advances in DNA sequencing. It included a review by Jay Schendure and Eriz Lieberman Aiden entitled “The Expanding Scope of DNA Sequencing [1],” in which the authors provided a great overview of current and future sequencing-based assay methods with an interesting technical twist. It also made for an opportunity to update a previous Finchtalk.

As DNA sequencing moved from determining the order of nucleotide bases in single genes to the factory style efforts of the first genomes, it was limited to measuring ensembles of molecules derived from single clones or PCR amplicons as composite sequences. Massively parallel sequencing changed the game because each molecule in a sample is sequenced independently. This discontinuous advance resulted in a massive increase in throughput that created a brief, yet significant, deviation in the price performance curve that would be predicted from Moore’s law. It also created a level of resolution that makes it possible to collect data from populations of sequences and see how they vary in a quantitative fashion making it possible to use DNA sequencing as a powerful assay platform. While this was quickly recognized [2], reducing ideas to practice would take a few more years.

Sequencing applications fall into three three main branches: De Novo, Functional Genomics, and Genetics (figure below). The De Novo, or Exploratory branch contains three subbranches: new genomes, meta-genomes, or meta-transcriptomes. Genetics or variation assays form another main branch of the tree. Genomic sequences are compared within and between populations, individuals, or tissue and cells with the goal predicting a phenotype from differences between sequences. Genetic assays can focus on single nucleotide variations, copy number changes or structural differences. Determining inherited epigenetic modifications is another form of genetic assay.

Understanding the relationship between genotype and phenotype, however, requires that we understand phenotype in sufficient detail. In order for this to happen, traditional analog measurements such as height, weight, blood pressure, and disease descriptions need to be replaced with quantitative measurements at the DNA, RNA, protein, metabolism, and other levels. Within each set of “omes” we need to understand molecular interactions and the how the environmental factors such as diet, chemicals, and microorganisms impact these interactions positively or negatively and through modification of the epigenome. Hence, the Functional Genomics branch is fastest growing.

New assays since 2010 are highlighted in color and underlined text.  See [1] for descriptions.
Functional Genomics experiments can be classified into five groups: Regulation, Epi-genomics, Expression, Deep Protein Mutagenesis, and Gene Disruption. Each group can be further divided into specific assay groups (DGE, RNA-Seq, small RNA, etc) that can be even further subdivided into specialized procedures (RNA-Seq with strandedness preserved). When experiments are refined and made reproducible, they become assays with sequence-based readouts.

In the paper, Shendure and Aiden describe 24 different assays. Citing an analogy to language where "Wilhelm von Humboldt described language as a system that makes ‘infi- nite use of finite means’: despite a relatively small number of words and combinatorial rules, it is possible to express an infinite range of ideas," the authors presented assay evolution as a assemblage of a small number of experimental designs. This model is not limited to language. In biochemistry a small number of protein domains and effector molecules are combined, and slightly modified, in different ways to create a diverse array of enzymes, receptors, transcription factors, and signaling cascades.

Subway map from [1]*. 
Shendure and Aiden go on show how the technical domains can be combined to form new kinds of assays using a subway framework, where one enters via a general approach (comparison, perturbation, or variation) and reaches the final sequencing destination. Stations along the way are specific techniques that are organized by experimental motifs including cell extraction, nucleic acid extraction, indirect targeting, exploiting proximity, biochemical transformation, and direct DNA or RNA targeting.

The review focused on the bench and made only brief reference to the informatics issues as part of the "rate limiters" of next-generation sequencing experiments.  It is important to note that each assay will have its own data analysis methodology. That may seem daunting. However, like the assays, the specialized informatics pipelines and other analyses can also be developed from a common set of building blocks. At Geospiza we are very familiar with these building blocks and how they can be assembled to analyze the data from many kinds of assays. As a result, the GeneSifter system is the most comprehensive in terms of its capabilities to support a large matrix of assays, analytical procedures, and species.  If you are considering adding next-generation sequencing to your research or your current informatics is limiting your ability to publish, check out GeneSifter.

1. Shendure, J., and Aiden, E. (2012). The expanding scope of DNA sequencing Nature Biotechnology, 30 (11), 1084-1094 DOI: 10.1038/nbt.2421

2. Kahvejian A, Quackenbush J, and Thompson JF (2008). What would you do if you could sequence everything? Nature biotechnology, 26 (10), 1125-33 PMID: 18846086

* Rights obtained from Rightslink number 3084971224414

Sunday, January 27, 2013

Sneak Peek: Identifying Mutations in Expressed Regions of Genomes Using NGS

Join us Wednesday, January 30th at 1 PM (EST), 10 AM (PST) to learn how to use NGS to identify mutations in expressed regions of genomes.

Abstract:

The pace at which genome references are being generated for plants and animal species is rapidly increasing with Next Generation Sequencing technologies. While this is a major step forward for researchers studying species that previously did not have sequenced genomes, it is only the beginning of the process toward defining the biology underlying the genome. As long as a reference is available, DNA variants can be readily identified on a genome wide scale, often producing lists of 100s of thousands or even millions of variants. Frequently these variants that occur in expressed genes are of the most interest; however, if annotation defining where genes exist within a genome is not available or poorly defined, identifying which mutations might affect protein coding may not be possible. To address this challenge we will describe a method whereby RNA-Seq can be readily used to identify transcriptionally active regions which creates transcript annotation for un-annotated or enhanced annotation for any organism. This annotation can then be used in conjunction with whole genome sequencing to annotate variants as to whether they fall within transcriptionally active regions thus facilitating the identification of mutations in larger repertoire of expressed regions of a genome.

Eric Olson, Ph.D., and  Hugh Arnold, Ph.D., from Geospiza will present. 

Thursday, January 17, 2013

Bio Databases 2013

ResearchBlogging.org  I seem to have committed to an annual ritual of summarizing the Nucleic Acids Research (NAR) Database Issue [1]. I do this because it is important to understand and emphasize the increasing role of data analysis in modern biology and remind us about the challenges that persist in turning data into knowledge.

Sometimes I hear individuals say they are building a database of all knowledge. To them I say good luck! The reality is new knowledge is developed from unique insights that are derived from specialized aggregations of information. Hence, as more data become available, through decreasing data collection costs, the number of resources and tools that are used to organize, analyze, and annotate data and information increases. Interestingly data cost decreases result from increased production due to technical improvements, which is an exponential function, whereas database growth is linear. Collecting data is the easy part.

How many are there?

Databases live in the wild and thus are hard to count.  Reading the introduction to the database issue one would think 88 new databases were added (cited), but if you compare the number being tracked by NAR in 2012 (1380) to 2013 (1512), you get 132. Moreover, databases tracked by NAR are contributed by authors.  Some don't bother with this.  For example, SeattleSNPs, home of the SeattleSeq Annotation and important Genome Variant Servers*, is not listed in NAR.  Nevertheless the NAR registry continues to increase by about 100 databases per year.


What's new?

Last year, I noted that the new databases did not reflect any discrenable pattern in terms of how the field of biology was changing. Rather the new databases reflect increasing specialization and complexity.  That trend continues, but this year Fernández-Suárez and Galperin note the emergence of new databases for studying human disease. Altogether eight databases were cited in the introduction. Several others are listed in a table highlighting the new new databases. While databases specializing in human genetics are not new, the past year saw an increased emphasis on understanding the relationship between genotype and phenotype as we advance our understanding of rare variation and population genetics.

As noted, many databases support human genomics research. If you visit the NAR Database Summary Category List and expand the list of databases under the Human Genes and Diseases list, you find four sub categories (General human Genetics, general polymorphism, Cancer gene, and Gene-, system-, or disease-specific databases) listing approximately 174 database. I say approximately because, as noted above, databases are hard to count. Curiously, just above Human Genes and Diseases is a category called Human and Vertebrate Genomes. Database are hard to classify too.

What's useful?

It is clear that the growing number of databases reflects an increasing level of specialization. Also likely is a high degree of redundancy. 10 microRNA databases (found by virtue of starting with "miR") cover general and specific topics including miRNAs that are predicted from sequence or literature, verified by experiment as existing or having a target, being possibly pathogenic, or existing in different organisms. It would be interesting to see which of these databases have the same data, but that is hard as some sites make all data available and some make their data searchable only.  In the former case, getting the data requires that it be put into a common format to make comparisons. Hence, access and interoperability issues persist.

Databases also persist. Fernández-Suárez and Galperin commented on efforts to curate the NAR collection. The annual attrition rare is less than 5% and greater than 90% of the databases are functional as determined by their response to webbots. Some have merged into other projects. What is not known is the quality of the information. In other words how are databases verified for accuracy or maintained to reflect our changing state of knowledge? As databases become increasing used in medical sequencing caveat emptor changes to caveat venditor and validation will be a critical component of design and maintenance. Perhaps future issues of the NAR database update will comment on these challenges.

Reference:
[1] Fernández-Suárez XM, and Galperin MY (2013). The 2013 Nucleic Acids Research Database Issue and the online Molecular Biology Database Collection. Nucleic acids research, 41 (D1) PMID: 23203983

Footnote:
* The SeattleSeq and Genome Variant Server links will break at the next update because the URL's contain the respective database version numbers.

Tuesday, December 4, 2012

Commonly Rare

ResearchBlogging.org
Rare is the new common. The final month of the year is always a good time to review progress and think about what's next.  In genetics, massively parallel next generation sequencing (NGS) technologies have been a dominating theme, and for good reason.

Unlike the previous high-throughput genetic analysis technologies (Sanger sequencing and microarrays), NGS allows us to explore genomes in far deeper ways and measure functional elements and gene expression in global ways.

What have we learned?

Distribution of rare and common variants. From [1] 
The ENCODE project has produced a picture where a much greater fraction of the genome may be involved in some functional role than previously understood [1]. However, a larger theme has been related to observing rare variation, and trying to understand its impact on human health and disease. Because the enzymes that replicate DNA and correct errors are not perfect, each time a genome is copied a small number of mutations are introduced, on average between 35-80. Since sperm are continuously produced, fathers contribute more mutations than mothers, and the number of new mutations increases with the father's age [2]. While the number per child, with respect to their father's contributed three-billion base genome, is tiny, rare diseases and intellectual disorders can result.

A consequence is that the exponentially growing human population has accumulated a very large number of rare genetic variants [3]. Many of these variants can be predicted to affect phenotype and many more may modify phenotypes in yet unknown ways [4,5].  We are also learning that variants generally fall into two categories. They are either common to all populations or confined to specific populations (figure). More importantly, for a given gene the number of rare variants can vastly outnumber of the number of previously known common variants.

Another consequence of the high abundance of rare variation is how it impacts the resources that are used to measure variation and map disease to genotypes.  For example, microarrays, which have been the primary tool of genome wide association studies utilize probes developed from a human reference genome sequence. When rare variants are factored in, many probes have several issues ranging from "hidden" variation within a probe to a probe simply not being able to measure a variant that is present. Linkage block size is also affected [6]. What this means it the best arrays going forward will be tuned to specific populations. It also means we need to devote more energy to developing refined reference resources, because the current tools do not adequately account for human diversity [6,7].

What's next?

Rare genetic variation has been understood for sometime. What's new is understanding just how extensive these variants are in the human population, which has resulted from the population recently rapidly expanding under very little selective pressure.  Hence, linking variation to heath and disease is the next big challenge and the cornerstone of personalized medicine, or as some would like precision medicine. Conquering this challenge will require detailed descriptions of phenotypes, in many cases at the molecular level. As the vast majority of variants, benign or pathogenic, lie outside of coding regions we will need to deeply understand how those functional elements, as initially defined by ENCODE, are affected by rare variation. We will also need to layer in epigenetic modifications.

For the next several years the picture will be complex.

References:

1. 1000 Genomes Project Consortium (2012). An integrated map of genetic variation from 1,092 human genomes. Nature, 491 (7422), 56-65 PMID: 23128226

[2] Kong, A., et. al. (2012). Rate of de novo mutations and the importance of father’s age to disease risk Nature, 488 (7412), 471-475 DOI: 10.1038/nature11396

[3] Keinan, A., and Clark, A. (2012). Recent Explosive Human Population Growth Has Resulted in an Excess of Rare Genetic Variants Science, 336 (6082), 740-743 DOI: 10.1126/science.1217283

[4] Tennessen, J., et. al. (2012). Evolution and Functional Impact of Rare Coding Variation from Deep Sequencing of Human Exomes Science, 337 (6090), 64-69 DOI: 10.1126/science.1219240

[5] Nelson, M., et. al. (2012). An Abundance of Rare Functional Variants in 202 Drug Target Genes Sequenced in 14,002 People Science, 337 (6090), 100-104 DOI: 10.1126/science.1217876

[6] Rosenfeld JA, Mason CE, and Smith TM (2012). Limitations of the human reference genome for personalized genomics. PloS one, 7 (7) PMID: 22811759

[7] Smith TM., and Porter SG. (2012) Genomic Inequality. The Scientist.



Sunday, August 5, 2012

Remembering Chris Abajian

Chris Abajian was a change catalyst.  Using a biochemical analogy, passion, creativity, and intellect were his catalytic triad. Together, with Joe Slagel, Chris, and I started Geospiza in 1997.  Sadly, Chris recently died in a hiking accident (7/30/12).  In remembrance, I'll share a few stories from our times together.

I met Chris during my Postdoc in Leroy Hood's laboratory in 1994.  This was the early days of the human genome project and we hired Chris, because in Lee's view we were going to build the best software if we had professional software engineers on the team. Lee was right.

Chris accepted our offer and from his first day, he made it clear this was not just a job where he could  apply his software development talents, it was an opportunity to have an impact.  And he did.

Sputnik

Chris used his passion, creativity and intellect to identify problems that needed to be solved, and then advocate creative solutions, passionately. One of his first programs was Sputnik - a tool that identifies microsatellite sequences.  Sputnik was inspired by a co-worker of ours Lee Rowen.  One day Chris observed Lee hunched over a ream of paper with printed DNA sequences one hand and a highlighter in the other. When he inquired as to what she was doing she responded "identifying microsatelites."

New to biology Chris asked what those were.  Lee explained that they are small repeating patterns of di, tri, tetra, or slightly longer sequences and that we are interested in them because they can be involved in disease and change gene regulation.  Chris quickly went to work. He talked to everyone around, learned that the repeated patterns were not always perfect and used this information to develop an algorithm and scoring table that could identify micro-satellite patterns with pretty good accuracy.  

Today a google search on "sputnik microsatellite" or "sputnik-microsatelite" yields ~191,000 or ~65,000 hits, respectively.  What's even more interesting is the number of papers, 12 and 15 years later, that compare different microsatellite algorithms to sputnik [1,2].  Not bad for a music major without any formal biology training!

Consed

After sputnik, Chris turned his attention to the next problem.  This was in the early days of Phred and Phrap (P. Green, still unpublished) and we had no way to work with DNA sequence assemblies in graphical user interface (GUI).   Not everyone was convinced we needed to build a whole new application. It would be a significant undertaking and other tools could be hacked to view Phrap assemblies.  This is when we learned that when Chris set out to do something, he was going to get it done and do so convincingly. To Chris, and some others, it was clear Phrap needed its own GUI, so he set to work, debated the points and got buy-in. In collaboration with David Gordon, Chris proceeded to build Consed.  Chris worked on the project for only a short time, but the work was a success. 17 years later, David has developed a large loyal user base and continues to develop new features for Consed [3].

The hunt for BRCA1

Chris and I worked closely on many software development projects, starting with data delivery for BRCA1. In 1994, we were asked to help "hunt" for the gene. In collaboration with Mary-Claire King, Francis Collins, Maynard Olson, and Lee Hood, we set out to find the BRCA1 gene. It had been previously localized to a large region of chromosome 17 by the King and Collins groups, and they had created a cosmid library of the region. With the high-throughput sequencing technology of 1994 we could include DNA sequencing in our strategy; one cosmid at a time.  So, our job in the lab was to get cosmid DNA clones from the King and Collins libraries, sequence them and make the data available to everyone -  simultaneously. How were we going to do that?

With web-technology

In 1994 the Mosiac web browser was new.  Chris suggested that we could post the sequences to website and send emails to the respective parties when the data was posted. Problem solved!

During this time, I was learning to program and developing automation systems. It was a no brainer and we set to work, Chris created a framework that I could use to create automation scripts. This was going to be a theme that would result in several more successful projects and lead to the next adventure. 

Geospiza

One day in April of 1997, Chris, Joe, and myself squeezed into the cab of Chris's small Toyota truck and headed to airport to interview with a new bioinformatics company called Pangea.  We were hired, but it was soon clear that we needed to do something different.  After a few rounds of passionate conversation we knew we were going to form a company, and we did.

Geospiza started in October of 1997 and while Chris was with us for only a short period of time, he made contributions that would last. Geospiza continues, now within PerkinElmer and there are probably still a few lines of his original code working within our LIMS system.

It's amazing to think about his accomplishments over the four years we spent together.  We enjoyed many good times discussing science, literature, and music. Chris will be missed.

References 
1. Kofler, R. (2007-07-01) SciRoKo: a new tool for whole genome microsatellite search and investigation. Bioinformatics, 7(4), 524-1685. DOI: 10.1093/bioinformatics/btm157

2. Leclercq, S. (2007) Detecting microsatellites within genomes: significant variation among algorithms. BMC Bioinformatics, 8(1), 125. DOI: 10.1186/1471-2105-8-125

3. David Gordon, Chris Abajian, and Phil Green. Consed: a graphical tool for sequence finishing. Genome Res. 1998. 8: 195-202

Thursday, July 12, 2012

Resources for Personalized Medicine Need Work

Yesterday (July 11, 2012), PLoS ONE published an article prepared by my colleagues and myself entitled "Limitations of the Reference Genome for Personalized Genomics."

This work, supported by Geospiza's SBIR targeting ways to improve mutation detection and annotation, explored some the resources and assumptions that are used to measure and understand sequence variation.  As we know, a key deliverable of the human genome project was to produce a high quality reference sequence that could be used to annotate genes, develop research tools like genotyping and microarray assays, and provide insights to guide software development. Projects like HapMap used these resources to provide additional understandings in terms of genetic linkage in populations.

Decreasing sequencing costs
Since those early projects, DNA sequencing costs have plummeted.  As a result, endeavors such as the 1000 Genomes Project (1KGP) and public contributions from Complete Genomics (CG) have dramatically increased the number of known sequence variants.  A question worth asking is how do these new data contribute to an understanding of the utility of current resources and assumptions that have guided genomics and genetics for the past six or seven years?

Number of variants by dbSNP build
To address the above question, we evaluated several assay and software tools that were based on the human genome reference sequence in the context of new data contributed by 1KGP and CG. We found a high frequency of confounding issues with microarrays, and many cases where invalid assumptions, encoded in bioinformatics programs, underestimate variability or possibly misidentify the functional effects of mutations. For example, 34% of published array-based GWAS studies for a variety of diseases utilize probes that contain undocumented variation or map to regions of previously unknown structural variation. Similarly, assumptions about the size of linkage disequillibrium decrease as the numbers of variants increase.

The significance of this work is that it documents what many are anecdotally experiencing. As we continue to learn about the contributing role of rare variation in human disease we need to fully understand how current resources can be used and work to resolve discrepancies in order to create an era of personalized medicine.

(2012). Limitations of the Human Reference Genome for Personalized Genomics, PLoS ONE,   DOI: 10.1371/journal.pone.0040294.t002

Tuesday, May 8, 2012

FinchTV on Lion

Okay, it took a while, but FinchTV, Geospiza's popular Sanger sequencing trace viewer is now available on Mac OS X Lion.

While the same great features are still great, the underlying code has been updated to run the application as a native intel binary, so it can be great into the future.

Features make FinchTV cool


In addition to all the basic things you'd expect a trace viewer to do like open AB1 or SCF files, view bases, electropherogram peaks, quality values, and reverse complement sequences, and dynamically scale data, FinchTV also lets you edit bases, print traces, and view detailed information about your trace file.  And, you can open files with a simple drag and drop action, as you'd expect for a modern desktop application.

That's not all, there's more

What really makes FinchTV stand out, is the ability to view your trace in a single-pane or multi-pane view. The latter view is ideal for visualizing the full data contained in the trace. In multi-pane view you can even change the horizontal or vertical scales.

FinchTV also integrates with NCBI's BLAST services.  In the application you can highlight a region of sequence and either use the edit menu or right click with your mouse to get the BLAST menu to choose between nucleotide (BLASTn), translated nucleotide (BLASTx), translated nucleotide + translated database (TBLASTx), or mega BLAST options.

Less obvious features include the ability to search for sub sequences in your DNA sequence and observe the raw data for a trace. Subsequence searching uses Perl style regular expressions and a Greedy algorithm so you can enter a search term and, each time you hit return, the next best match is identified including subsequences within subsequences. For example, the regular expression - ATG((?!TAG|TAA|TGA)...)+(TAG|TAA|TGA) - will find open reading frames. If one reading frame is contained within another, the longer reading frame is found first.

Another less known feature is the raw data view.  The raw data view is essential for those work with sequencing instruments.  Many do not realize the standard electropherogram trace image is processed data. That is, a mathematical matrix computation is applied to normalize signals, subtract background fluorescence, and correct natural mobility shifts in the data. The result is an easy to interpret view of the data. However, if you need to troubleshoot your sequencer and see the real details you'll want to view the rawest data possible. FinchTV is the only available trace viewer, outside of ABI's software that provides this capability.

To obtain FinchTV please visit http://www.geospiza.com/Products/finchtv.shtml




Sunday, April 22, 2012

Sneak Peak: A Practical Approach to Detecting Nucleotide Variants in NGS Data


Join us Thursday, May 3, 2012 9:00 am (Pacific Time) for a webinar on analyzing DNA sequencing data with hundreds of thousands to millions of nucleotide variants.

Description:
This webinar discusses DNA variant detection using Next Generation Sequencing for targeted and exome resequencing applications as well as, whole transcriptome sequencing. The presentation includes an overview of each application and its specific data analysis needs and challenges with a particular emphasis on variant detection methods and approaches for individual samples as well as multi-sample comparisons. For in depth comparisons of variant detection methods, Geospiza’s cloud-based GeneSifter® Analysis Edition software will be used to assess sample data from NCBI’s GEO and SRA.

For more information, please visit the registration page.