Thursday, January 5, 2012

Bio Databases 2012


ResearchBlogging.orgLet's get 2012 started with an update on the growth of biological databases.  About this time last year, I summarized Nucleic Acids Research's (NAR) annual database issue where authors submit papers describing updates to existing databases and present new databases. In that post, I predicted that we would see between 60 and 120 new databases in 2011.  This year's update included 92 new databases [1].

How have things changed?
Overall the number of databases being tracked by NAR has grown from 1330 in 2011 to 1380 in 2012 (left hand figure). As 92 new databases were added, 42 must have been dropped. Interestingly, when one views the new database list page only 90 databases are listed.  I not only counted this using my fingers and toes - more than twice, but also copied the list into Excel to verify my original count. I wonder what the two, but unaccounted, new databases are? Never mind the 42 that disappeared, those are not listed. 

What are the new databases?

Reviewing the list of the 90 new databases, does not reveal any clear patterns or trends regarding changes in biology. Instead, the list reflects increasing complexity and specialization. Some databases tackle new kinds of data, many appear to be refinements of existing databases, and others contain highly specific information. For example, the UMD-BRCA1/ BRCA2 database contains information about BRCA1 and BRCA2 mutations detected in France.  A few databases, such as the BitterDB: a database of bitter taste molecules and receptors,  and Newt-omics: data on the red spotted newt (Notophthalmus viridescens), and IDEAL: Intrinsically Disordered proteins with Extensive Annotations and Literature, caught my attention because of their interesting descriptions. I especially like Disease Ontology: Ontology for a variety of human diseases. I wonder which diseases are their favorites?

Also of interest is the growth of wiki's as databases. This year, 10 of the new databases are wiki's where the community is invited to make contributions to the resource in a similar fashion to wikipedia.  I should say almost 10 of the databases are wiki's, one, SeqAnswers is actually a forum, that's technically not a wiki, of course some might argue that a wiki might not be a database.

So, what does this mean?

The growing list of databases is an interesting way to think about biology, DNA, and proteins. Simply examining the list is instructional and one can think of ways that mining several resources together could create new insights.  However, therein lies the challenge.  Many of these resources are not designed to interoperate and it is not clear how long they will last or be updated. This final point was made at the close of the introduction article. In a section entitled "Sustainability of bioinformatics databases," the authors discussed the past year's controversy surrounding NCBI's SRA database, previous challenges with Swiss-Prot, and the current instability with the KEGG and TAIR databases.  They cite a proposal to centralize more resources [2].

But in the end biological databases really do model biology and follow the principals of evolution. Perhaps it is apropos that new ones emerge through speciation, while others go extinct.

1. Galperin, M., & Fernandez-Suarez, X. (2011). The 2012 Nucleic Acids Research Database Issue and the online Molecular Biology Database Collection Nucleic Acids Research, 40 (D1) DOI: 10.1093/nar/gkr1196

2. Parkhill J, Birney E, Kersey P. Genomic information infrastructure after the deluge. Genome Biol. 2010;11:402


Tuesday, November 8, 2011

BioData at #SC11

Next week, Nov 12-18, Super Computing comes to Seattle. On Wed, Nov 15 at 12:15-1:15 pm, @finchtalk (me) will host a Birds-of-a-Feater session on "Technologies for Managing BioData" in room TCC305.

I'll kick off the session by sharing stories from Geospiza's work experiences and the work of others. If you have a story to share please bring it. The session will provide an open platform. We plan to cover relational databases, HDF5 technologies, and NoSQL. If you want to join in because you are interested in learning, the abstract below will give you an idea of what will be discussed.

Abstract:

DNA sequencing and related technologies are producing tremendous volumes of data. The raw data from these instruments needs to be reduced through alignment or assembly into forms that can be further processed to yield scientifically or clinically actionable information. The entire data workflow process requires multiple programs and information resources. Standard formats and software tools that meet high performance computing requirements are lacking, but technical approaches are emerging. In this BoF, options such as BAM, BioHDF, VCF and other formats, and corresponding tools, will be reviewed for their utility in meeting a broad set of requirements. The goal of the BoF is look beyond DNA sequencing and discuss the requirements for data management technologies that can integrate sequence data with data collected from other platforms such as quantitative PCR, mass spectrometry, and imaging systems. We will also explore the technical requirements for working with data from large numbers of samples.

Thursday, October 13, 2011

Personalities of Personal Genomes


"People say they want their genetic information, but they don’t." "The speaker's views of data return are frankly repugnant." These were some of the [paraphrased] comments and tweets expressed during Cold Spring Harbor's fourth annual conference entitled "Personal Genomes" held Sep 30 - Oct 2, 2011. The focus of which was to explore the latest technologies and approaches for sequencing genomes, exomes, and transcriptomes in the context of how genome science is, and will be, impacting clinical care. 

The future may be close than we think

In previous years, the concept of personal genome sequencing as a way to influence medical treatment was a vision. Last year, the reality of the vision was evident through a limited number of examples. This year, several new examples were presented along with the establishment of institutional programs for genomic-based medicine. The driver being the continuing decreases in data collection costs combined with corresponding access to increasing amounts of data. According to Richard Gibbs (Baylor College of Medicine) we will have close to 5000 genomes completely sequenced by the end of this year and by the end of 2012, 30,000 complete genome sequences are expected.

The growth of genome sequencing is now significant  enough that leading institutions are also beginning to establish guidelines for genomics-based medicine. Hence, an ethics panel discussion was held during the conference. The conversation about how DNA sequence data may be used has been an integral discussion since the beginning of the Genome Project. Indeed James Watson shared his lament for having to fund ethics research and directly asked the panel if they have done any good. There was a general consensus, from the panel, and audience members who have had their genomes sequenced, that ethics funding has helped by establishing genetic counseling and eduction practices.  

However, as pointed out by some audience members, this ethics panel, like many others, focused too heavily on the risks for individuals and society having their genomic data. In my view, the discussion would have been more interesting and balanced if the panel included the individuals who are working outside of institutions with new approaches  for understanding health. Organizations like 23andMe, Patients LIke Me, or the Genetic Alliance bring a very different and valuable perspective to the conversation.

Ethics was a fraction of the conference. The remaining talks at were organized into six sessions that covered personal cancer genomics, medically actionable genomics, personal genomes, rare diseases, and clinical implementations of personal genomics. The key messages from these presentations and posters was that, while genomics-based medical approaches have demonstrated success, much more research needs to be done before such approaches are mainstream.  

For example, in the case of cancer genomics, whole genome sequences from tumor and normal cells can give a picture of point mutations and structural rearrangements, but these data need to be accompanied by exome sequences to get the high read depth needed to accurately detect the low levels of rare mutations that may be disregulating cell growth or conferring resistance to treatment. Yet, the resulting profiles of variants are still inadequate to fully understand the functional consequences of the mutations. For this, transcriptome profiling is needed, and that is just the start.  

Once the data are collected they need to be processed in different ways, filtered, and compared within and between samples. Information from many specialized databases will be used in conjunction with statistical analyses to develop insights that can be validated through additional assays and measurements.  Finally, a lab seeking to do this work, and return results back to patients, will also need to be certified, minimally by CLIA standards.  For many groups this is significant undertaking, and good partners with experience and strong capabilities like PerkinElmer will be needed. 

Further Reading

Nature Coverage, Oct 6 issue:
Genomes on prescription
Other news and information:


Wednesday, August 10, 2011

Stitching Protein-Protien Interactions via DNA Sequencing

ResearchBlogging.org Stitch-Seq, one of the newest editions to the Next Generation DNA Sequencing (NGS) was presented in June's Nature Methods.

Back in 2008, when groups were realizing the power of NGS technologies, I entitled a post "Next Gen Sequencing is not Sequencing DNA" to make the point that massively parallel ultra-high throughput DNA sequencing could be used to for quantitative assays that can measure transcriptome expression, protein-DNA interactions, methylation patterns, and more. Stitch-seq can now be added to a growing list of assays that include RNA-Seq, DNAse-Seq, ChIP-Seq, or HITS-CLIP, and others.

Stitch-seq explores the interactome, a term used to describe how the molecules of a cell interact in networks to carryout life's biochemical activities. Understanding how these networks are controlled through genetics and environmental stimuli is critical in discovering biomarkers that can be used to stratify disease and target highly specific therapies. However, the interactome is complex; studying it requires that interactions can be identified at high scale.

Many interactome studies focus on proteins. Traditional approaches involve specially constructed gene reporter systems. For example, in the two-hybrid approach, a portion of a protein encoding gene is combined with a gene fragment containing a DNA binding domain of a transcription factor (bait). In another construct a different protein encoding region is combined with the RNA polymerase binding domain fragment of the same transcription factor (prey).

When the DNA constructs are expressed, interactions can be measure by gene expression. If the protein attached to the bait interacts with the protein attached to the prey, transcription is initiated at the reporter gene.  When reporter genes confer growth on selective media, interacting protein encoding segments can be identified by isolating the DNA from growing cells and sequencing the DNA constructs.

Therein lies the rub

Until now, interactome studies combined high-throughtput assays systems with low-throughput characterization systems that PCR amplified the individual constructs and characterized the DNA by Sanger sequencing. Yu and colleagues overcame this problem by devising a new strategy that put potential interacting domains on common DNA fragments, via "stitch-PCR" to prepare libraries that can easily be sequenced by NGS methods.  Using this method the team was able to increase overall assay throughput by 42% and measure 1000s of interactions.

While still low-throughput relative to the kinds of numbers were used to on NGS, increasing the throughput of protein interaction assays is an important step toward making systems biology experiments more scalable. It also adds another Seq to our growing collection of Assay-Seq methods.

Yu, H., Tardivo, L., Tam, S., Weiner, E., Gebreab, F., Fan, C., Svrzikapa, N., Hirozane-Kishikawa, T., Rietman, E., Yang, X., Sahalie, J., Salehi-Ashtiani, K., Hao, T., Cusick, M., Hill, D., Roth, F., Braun, P., & Vidal, M. (2011). Next-generation sequencing to generate interactome datasets Nature Methods, 8 (6), 478-480 DOI: 10.1038/nmeth.1597

Friday, July 8, 2011

It's not just science and technology

PerkinElmer is committed to improving human and environmental health. Sometimes that extends beyond delivering commercial products and services to how we participate in our community.

Geospiza was proud to help the local community last Tuesday (7/5) by volunteering, through PerkinElmer's Corporate Social Responsibility program, to clean up the area around Lake Union after the Seattle's Fourth of July Celebration.

It's also great way to post a couple of pictures of some of the Geospiza team.

Pictured (l-r): Karen Friery, Sarah Lauer, Stephanie Tatem Murphy, Darrell Reising, Jackie Wright and    Jeff Kozlowski. 

Pictured (l-r): Darrell Reising, Jim Hancock, Sarah Lauer, Stephanie Tatem Murphy, Todd Smith, Jackie Wright and Jeff Kozlowski

About the activity

After thousands celebrate the Fourth of July and the watch the fireworks, Seattle’s Lake Union can be more trampled fields of chip bags and shining mounds of beer cans than amber waves of grain and purple mountains majesty.

As a member of the biotech community based around Lake Union, Geospiza donned gardening gloves and joined our neighbors for the 5th of July Lake Union Cleanup, a post-fireworks cleaning frenzy hosted by Starbucks, Puget Sound Keeper Alliance and Seattle Public Utilities.

Last year’s inaugural event had more than 700 volunteers collecting 1.15 tons of garbage – one-third of which was recycled. The cleaning event grew out of the community-based efforts to save the fireworks. When the 2010 show was cancelled due to lack of corporate sponsorship, the community launched a fund drive which raised enough money in less than 24hrs.

From April through the end of July, every site across PerkinElmer, including Geospiza, participates in one “For The Better Day” event to fulfill our company mission to improve human and environmental health while providing employees with an opportunity to build stronger teams and communities where we work, live and play.

Friday, June 10, 2011

Sneak Peak: NGS Resequencing Applications: Part I – Detecting DNA Variants

Join us next Wed. June 15 for a webinar on resequencing applications.

Description
This webinar will focus on DNA variant detection using Next Generation Sequencing for the applications of targeted and exome resequencing as well as, whole transcriptome sequencing. The presentation will include an overview of each application and its specific data analysis needs and challenges. Topics covered will include Secondary Analysis (alignments, reference choices, variant detection) with a particular emphasis on DNA variant detection as well as multi-sample comparisons. For in depth comparisons of variant detection methods, Geospiza’s cloud-based GeneSifter Analysis Edition software will be used to assess sample data from NCBI’s GEO and SRA. The webinar will also include a short presentation on how these tools can be deployed for both individual researchers as well as through Geospiza’s Partner Program for NGS sequencing service providers.

Details:
Date and time: Wednesday, June 15, 2011 10:00 am
Pacific Daylight Time (San Francisco, GMT-07:00)
Wednesday, June 15, 2011 1:00 pm 
Eastern Daylight Time (New York, GMT-04:00)
Wednesday, June 15, 2011 6:00 pm
GMT Summer Time (London, GMT+01:00)
Duration: 1 hour

Tuesday, June 7, 2011

DOE's 2011 Sequencing, Finishing, Analysis in the Future Meeting

Cactus at Bandelier
National Monument 
Last week, June 1-3, the Department of Energy held their annual Sequencing, Finishing, Analysis in the Future (SFAF) meeting in Santa Fe, New Mexico.  SFAF, also sponsored b the Joint Genome Institute, and Los Alamos National Laboratory and was attended by individuals from the major genome centers, commercial organizations, and smaller labs.

In addition to standard presentations and panel discussions from the genome centers and sequencing vendors (Life Technologies, Illumina, Roche 454, and Pacific Biosciences), and commercial tech talks, this year's meeting included a workshop on hybrid sequence assembly (mixing Illumina and 454 data, or Illumina and PacBio data). I also presented recent work on how 1000 Genomes and Complete Genomics data are changing our thinking about genetics (abstract below).

John McPherson from the Ontario Cancer Research Institute (OICR, a Geospiza client) gave the kickoff keynote. His talk focused on challenges in cancer sequencing. One of those being that DNA sequencing costs are now predominated by instrument maintenance, sample acquisition, preparation, and informatics, which are never included in the $1000 genome conversation. OICR is now producing 17 trillion bases per month and as they, and others, learn about cancer's complexity, the idea of finding single biochemical targets for magic bullet treatments is becoming less likely.

McPherson also discussed how OICR is getting involved clinical cancer sequencing. Because cancer is a genetic disease, measuring somatic mutations and copy number variations will be best for developing prognostic biomarkers. However, measuring such biomarkers in patients in order to calibrate treatments requires a fast turnaround time between tissue biopsy, sequence data collection, and analysis. Hence, McPherson sees IonTorrent and PacBio as the best platforms for future assays. McPherson closed his presentation stating that data integration is the grand challenge.  We're on it!

The remaining talks explored several aspects of DNA sequencing ranging from high throughput single cell sample preparation, to sequence alignment and de novo sequence assembly, to education and interesting biology. I especially liked Dan Distal's (New England Biolabs) presentation on the wood eating microbiome of shipworms. I learned that shipworms are actually little clams that use their shells as drills to harvest the wood. Understanding how the bacteria eat wood is important because we may be able to harness this ability for future energy production.

Finally, there was my presentation for which I've included the abstract.

What's a referenceable reference?

The goal behind investing time and money into finishing genomes to high levels of completeness and accuracy is that they will serve as a reference sequences for future research. Reference data are used as a standard to measure sequence variation, genomic structure, and study gene expression in microarray and DNA sequencing assays. The depth and quality of information that can be gained from such analyses is a direct function of the quality of the reference sequence and level of annotation. However, finishing genomes is expensive, arduous work. Moreover, in the light of what we are learning about genome and species complexity, it is worthwhile asking the question whether a single reference sequence is the best standard of comparison in genomics studies.

The human genome reference, for example, is well characterized, annotated, and represents a considerable investment. Despite these efforts, it is well understood that many gaps exist in even the most recent versions (hg19, build 37) [1], and many groups still use the previous version (hg18, build 36). Additionally, data emerging from the 1000 Genomes Project, Complete Genomics, and others have demonstrated that the variation between individual genomes is far greater than previously thought. This extreme variability has implications for genotyping microarrays, deep sequencing analysis, and other methods that rely on a single reference genome. Hence, we have analyzed several commonly used genomics tools that are based on the concept of a standard reference sequence, and have found that their underlying assumptions are incorrect. In light of these results, the time has come to question the utility and universality of single genome reference sequences and evaluate how to best understand and interpret genomics data in ways that take a high level of variability into account.

Todd Smith(1), Jeffrey Rosenfeld(2), Christopher Mason(3). (1) Geospiza Inc. Seattle, WA 98119, USA (2) Sackler Institute for Comparative Genomics, American Museum of Natural History, New York, NY 10024, USA (3) Weill Cornell Medical College, New York, NY 10021, USA

Kidd JM, Sampas N, Antonacci F, Graves T, Fulton R, Hayden HS, Alkan C, Malig M, Ventura M, Giannuzzi G, Kallicki J, Anderson P, Tsalenko A, Yamada NA, Tsang P, Kaul R, Wilson RK, Bruhn L, & Eichler EE (2010). Characterization of missing human genome sequences and copy-number polymorphic insertions. Nature methods, 7 (5), 365-71 PMID: 20440878

You can obtain abstracts for all of the presentations at the SFAF website.

Friday, May 20, 2011

21st Century Medicine: A Question of Ps

Last Sunday and Monday (5/15, 5/16/11) the Institute for Systems Biology (ISB) held their annual symposium. This year was the 10th annual and focused on "Systems Biology and P4 Medicine."

For those new to P4 medicine, the Ps stand for Personalized, Predictive, Preventative, and Participatory. P4 medicine is about changing our current disease oriented, reactive, approaches to those that prevent disease by increasing the predictive power of diagnostics. Because we are all different, future diagnostics need to be tailored to each individual, which also means individuals need to be more aware of their health and proactively participate in their health care. The vision of P4 medicine is that it will not only dramatically improve the quality of health care, it will significantly decrease health care costs. Hence, some folks add additional Ps to include payment and policy.

P4 medicine is an ambitious goal. In Lee Hood's closing notes he noted four significant challenges that need to be overcome to make P4 medicine a practical reality:

  1. IT challenges. In addition to working on how to transform datasets containing billions of measurements into actionable information, we need to integrate high dimensional data a wide variety of measurement systems. Reduced data will need to be presented in medical records that can be easily accessed and understood by health care providers and participants.
  2. Education. Students, scientists, doctors, individuals, and policy makers need to learn and develop an understanding of how the networks of interacting proteins and biochemicals that make us healthy or sick are regulated by our genomes and respond to environmental factors. 
  3. Big vs Small Science. Funding agencies are concerned with how to best support the research needed to create the kinds of technologies and approaches that will unlock biology's complexity to develop future diagnostics and efficacious therapies. Large-scale projects conducted over the past 10 years have made it clear that biology is extremely complex. Deciphering this complexity requires that we integrate production-orientated data collection approaches, that develop a data infrastructure, with focused research projects, run by domain experts, that explore specific ideas.  The challenge is balancing big and small science to achieve high impact goals. 
  4. Families. Understanding the genetic basis of health and disease requires the research be conducted on samples derived from families rather than randomized populations. Many families are needed to develop critical insights. However, in the U.S. our IRB (Institutional Review Boards) are considered a hinderance to enrolling individuals. 
Through the day and half conference numerous presentations explored different aspects of the above challenges. Walter Jessen at Biomarker Commons has created excellent summaries of the first and second day's presentations. For those who like raw data, the #ISB2011P4 hashtag can be used to get the symposium's tweets.