Tuesday, November 8, 2011

BioData at #SC11

Next week, Nov 12-18, Super Computing comes to Seattle. On Wed, Nov 15 at 12:15-1:15 pm, @finchtalk (me) will host a Birds-of-a-Feater session on "Technologies for Managing BioData" in room TCC305.

I'll kick off the session by sharing stories from Geospiza's work experiences and the work of others. If you have a story to share please bring it. The session will provide an open platform. We plan to cover relational databases, HDF5 technologies, and NoSQL. If you want to join in because you are interested in learning, the abstract below will give you an idea of what will be discussed.

Abstract:

DNA sequencing and related technologies are producing tremendous volumes of data. The raw data from these instruments needs to be reduced through alignment or assembly into forms that can be further processed to yield scientifically or clinically actionable information. The entire data workflow process requires multiple programs and information resources. Standard formats and software tools that meet high performance computing requirements are lacking, but technical approaches are emerging. In this BoF, options such as BAM, BioHDF, VCF and other formats, and corresponding tools, will be reviewed for their utility in meeting a broad set of requirements. The goal of the BoF is look beyond DNA sequencing and discuss the requirements for data management technologies that can integrate sequence data with data collected from other platforms such as quantitative PCR, mass spectrometry, and imaging systems. We will also explore the technical requirements for working with data from large numbers of samples.

Thursday, October 13, 2011

Personalities of Personal Genomes


"People say they want their genetic information, but they don’t." "The speaker's views of data return are frankly repugnant." These were some of the [paraphrased] comments and tweets expressed during Cold Spring Harbor's fourth annual conference entitled "Personal Genomes" held Sep 30 - Oct 2, 2011. The focus of which was to explore the latest technologies and approaches for sequencing genomes, exomes, and transcriptomes in the context of how genome science is, and will be, impacting clinical care. 

The future may be close than we think

In previous years, the concept of personal genome sequencing as a way to influence medical treatment was a vision. Last year, the reality of the vision was evident through a limited number of examples. This year, several new examples were presented along with the establishment of institutional programs for genomic-based medicine. The driver being the continuing decreases in data collection costs combined with corresponding access to increasing amounts of data. According to Richard Gibbs (Baylor College of Medicine) we will have close to 5000 genomes completely sequenced by the end of this year and by the end of 2012, 30,000 complete genome sequences are expected.

The growth of genome sequencing is now significant  enough that leading institutions are also beginning to establish guidelines for genomics-based medicine. Hence, an ethics panel discussion was held during the conference. The conversation about how DNA sequence data may be used has been an integral discussion since the beginning of the Genome Project. Indeed James Watson shared his lament for having to fund ethics research and directly asked the panel if they have done any good. There was a general consensus, from the panel, and audience members who have had their genomes sequenced, that ethics funding has helped by establishing genetic counseling and eduction practices.  

However, as pointed out by some audience members, this ethics panel, like many others, focused too heavily on the risks for individuals and society having their genomic data. In my view, the discussion would have been more interesting and balanced if the panel included the individuals who are working outside of institutions with new approaches  for understanding health. Organizations like 23andMe, Patients LIke Me, or the Genetic Alliance bring a very different and valuable perspective to the conversation.

Ethics was a fraction of the conference. The remaining talks at were organized into six sessions that covered personal cancer genomics, medically actionable genomics, personal genomes, rare diseases, and clinical implementations of personal genomics. The key messages from these presentations and posters was that, while genomics-based medical approaches have demonstrated success, much more research needs to be done before such approaches are mainstream.  

For example, in the case of cancer genomics, whole genome sequences from tumor and normal cells can give a picture of point mutations and structural rearrangements, but these data need to be accompanied by exome sequences to get the high read depth needed to accurately detect the low levels of rare mutations that may be disregulating cell growth or conferring resistance to treatment. Yet, the resulting profiles of variants are still inadequate to fully understand the functional consequences of the mutations. For this, transcriptome profiling is needed, and that is just the start.  

Once the data are collected they need to be processed in different ways, filtered, and compared within and between samples. Information from many specialized databases will be used in conjunction with statistical analyses to develop insights that can be validated through additional assays and measurements.  Finally, a lab seeking to do this work, and return results back to patients, will also need to be certified, minimally by CLIA standards.  For many groups this is significant undertaking, and good partners with experience and strong capabilities like PerkinElmer will be needed. 

Further Reading

Nature Coverage, Oct 6 issue:
Genomes on prescription
Other news and information:


Wednesday, August 10, 2011

Stitching Protein-Protien Interactions via DNA Sequencing

ResearchBlogging.org Stitch-Seq, one of the newest editions to the Next Generation DNA Sequencing (NGS) was presented in June's Nature Methods.

Back in 2008, when groups were realizing the power of NGS technologies, I entitled a post "Next Gen Sequencing is not Sequencing DNA" to make the point that massively parallel ultra-high throughput DNA sequencing could be used to for quantitative assays that can measure transcriptome expression, protein-DNA interactions, methylation patterns, and more. Stitch-seq can now be added to a growing list of assays that include RNA-Seq, DNAse-Seq, ChIP-Seq, or HITS-CLIP, and others.

Stitch-seq explores the interactome, a term used to describe how the molecules of a cell interact in networks to carryout life's biochemical activities. Understanding how these networks are controlled through genetics and environmental stimuli is critical in discovering biomarkers that can be used to stratify disease and target highly specific therapies. However, the interactome is complex; studying it requires that interactions can be identified at high scale.

Many interactome studies focus on proteins. Traditional approaches involve specially constructed gene reporter systems. For example, in the two-hybrid approach, a portion of a protein encoding gene is combined with a gene fragment containing a DNA binding domain of a transcription factor (bait). In another construct a different protein encoding region is combined with the RNA polymerase binding domain fragment of the same transcription factor (prey).

When the DNA constructs are expressed, interactions can be measure by gene expression. If the protein attached to the bait interacts with the protein attached to the prey, transcription is initiated at the reporter gene.  When reporter genes confer growth on selective media, interacting protein encoding segments can be identified by isolating the DNA from growing cells and sequencing the DNA constructs.

Therein lies the rub

Until now, interactome studies combined high-throughtput assays systems with low-throughput characterization systems that PCR amplified the individual constructs and characterized the DNA by Sanger sequencing. Yu and colleagues overcame this problem by devising a new strategy that put potential interacting domains on common DNA fragments, via "stitch-PCR" to prepare libraries that can easily be sequenced by NGS methods.  Using this method the team was able to increase overall assay throughput by 42% and measure 1000s of interactions.

While still low-throughput relative to the kinds of numbers were used to on NGS, increasing the throughput of protein interaction assays is an important step toward making systems biology experiments more scalable. It also adds another Seq to our growing collection of Assay-Seq methods.

Yu, H., Tardivo, L., Tam, S., Weiner, E., Gebreab, F., Fan, C., Svrzikapa, N., Hirozane-Kishikawa, T., Rietman, E., Yang, X., Sahalie, J., Salehi-Ashtiani, K., Hao, T., Cusick, M., Hill, D., Roth, F., Braun, P., & Vidal, M. (2011). Next-generation sequencing to generate interactome datasets Nature Methods, 8 (6), 478-480 DOI: 10.1038/nmeth.1597

Friday, July 8, 2011

It's not just science and technology

PerkinElmer is committed to improving human and environmental health. Sometimes that extends beyond delivering commercial products and services to how we participate in our community.

Geospiza was proud to help the local community last Tuesday (7/5) by volunteering, through PerkinElmer's Corporate Social Responsibility program, to clean up the area around Lake Union after the Seattle's Fourth of July Celebration.

It's also great way to post a couple of pictures of some of the Geospiza team.

Pictured (l-r): Karen Friery, Sarah Lauer, Stephanie Tatem Murphy, Darrell Reising, Jackie Wright and    Jeff Kozlowski. 

Pictured (l-r): Darrell Reising, Jim Hancock, Sarah Lauer, Stephanie Tatem Murphy, Todd Smith, Jackie Wright and Jeff Kozlowski

About the activity

After thousands celebrate the Fourth of July and the watch the fireworks, Seattle’s Lake Union can be more trampled fields of chip bags and shining mounds of beer cans than amber waves of grain and purple mountains majesty.

As a member of the biotech community based around Lake Union, Geospiza donned gardening gloves and joined our neighbors for the 5th of July Lake Union Cleanup, a post-fireworks cleaning frenzy hosted by Starbucks, Puget Sound Keeper Alliance and Seattle Public Utilities.

Last year’s inaugural event had more than 700 volunteers collecting 1.15 tons of garbage – one-third of which was recycled. The cleaning event grew out of the community-based efforts to save the fireworks. When the 2010 show was cancelled due to lack of corporate sponsorship, the community launched a fund drive which raised enough money in less than 24hrs.

From April through the end of July, every site across PerkinElmer, including Geospiza, participates in one “For The Better Day” event to fulfill our company mission to improve human and environmental health while providing employees with an opportunity to build stronger teams and communities where we work, live and play.

Friday, June 10, 2011

Sneak Peak: NGS Resequencing Applications: Part I – Detecting DNA Variants

Join us next Wed. June 15 for a webinar on resequencing applications.

Description
This webinar will focus on DNA variant detection using Next Generation Sequencing for the applications of targeted and exome resequencing as well as, whole transcriptome sequencing. The presentation will include an overview of each application and its specific data analysis needs and challenges. Topics covered will include Secondary Analysis (alignments, reference choices, variant detection) with a particular emphasis on DNA variant detection as well as multi-sample comparisons. For in depth comparisons of variant detection methods, Geospiza’s cloud-based GeneSifter Analysis Edition software will be used to assess sample data from NCBI’s GEO and SRA. The webinar will also include a short presentation on how these tools can be deployed for both individual researchers as well as through Geospiza’s Partner Program for NGS sequencing service providers.

Details:
Date and time: Wednesday, June 15, 2011 10:00 am
Pacific Daylight Time (San Francisco, GMT-07:00)
Wednesday, June 15, 2011 1:00 pm 
Eastern Daylight Time (New York, GMT-04:00)
Wednesday, June 15, 2011 6:00 pm
GMT Summer Time (London, GMT+01:00)
Duration: 1 hour

Tuesday, June 7, 2011

DOE's 2011 Sequencing, Finishing, Analysis in the Future Meeting

Cactus at Bandelier
National Monument 
Last week, June 1-3, the Department of Energy held their annual Sequencing, Finishing, Analysis in the Future (SFAF) meeting in Santa Fe, New Mexico.  SFAF, also sponsored b the Joint Genome Institute, and Los Alamos National Laboratory and was attended by individuals from the major genome centers, commercial organizations, and smaller labs.

In addition to standard presentations and panel discussions from the genome centers and sequencing vendors (Life Technologies, Illumina, Roche 454, and Pacific Biosciences), and commercial tech talks, this year's meeting included a workshop on hybrid sequence assembly (mixing Illumina and 454 data, or Illumina and PacBio data). I also presented recent work on how 1000 Genomes and Complete Genomics data are changing our thinking about genetics (abstract below).

John McPherson from the Ontario Cancer Research Institute (OICR, a Geospiza client) gave the kickoff keynote. His talk focused on challenges in cancer sequencing. One of those being that DNA sequencing costs are now predominated by instrument maintenance, sample acquisition, preparation, and informatics, which are never included in the $1000 genome conversation. OICR is now producing 17 trillion bases per month and as they, and others, learn about cancer's complexity, the idea of finding single biochemical targets for magic bullet treatments is becoming less likely.

McPherson also discussed how OICR is getting involved clinical cancer sequencing. Because cancer is a genetic disease, measuring somatic mutations and copy number variations will be best for developing prognostic biomarkers. However, measuring such biomarkers in patients in order to calibrate treatments requires a fast turnaround time between tissue biopsy, sequence data collection, and analysis. Hence, McPherson sees IonTorrent and PacBio as the best platforms for future assays. McPherson closed his presentation stating that data integration is the grand challenge.  We're on it!

The remaining talks explored several aspects of DNA sequencing ranging from high throughput single cell sample preparation, to sequence alignment and de novo sequence assembly, to education and interesting biology. I especially liked Dan Distal's (New England Biolabs) presentation on the wood eating microbiome of shipworms. I learned that shipworms are actually little clams that use their shells as drills to harvest the wood. Understanding how the bacteria eat wood is important because we may be able to harness this ability for future energy production.

Finally, there was my presentation for which I've included the abstract.

What's a referenceable reference?

The goal behind investing time and money into finishing genomes to high levels of completeness and accuracy is that they will serve as a reference sequences for future research. Reference data are used as a standard to measure sequence variation, genomic structure, and study gene expression in microarray and DNA sequencing assays. The depth and quality of information that can be gained from such analyses is a direct function of the quality of the reference sequence and level of annotation. However, finishing genomes is expensive, arduous work. Moreover, in the light of what we are learning about genome and species complexity, it is worthwhile asking the question whether a single reference sequence is the best standard of comparison in genomics studies.

The human genome reference, for example, is well characterized, annotated, and represents a considerable investment. Despite these efforts, it is well understood that many gaps exist in even the most recent versions (hg19, build 37) [1], and many groups still use the previous version (hg18, build 36). Additionally, data emerging from the 1000 Genomes Project, Complete Genomics, and others have demonstrated that the variation between individual genomes is far greater than previously thought. This extreme variability has implications for genotyping microarrays, deep sequencing analysis, and other methods that rely on a single reference genome. Hence, we have analyzed several commonly used genomics tools that are based on the concept of a standard reference sequence, and have found that their underlying assumptions are incorrect. In light of these results, the time has come to question the utility and universality of single genome reference sequences and evaluate how to best understand and interpret genomics data in ways that take a high level of variability into account.

Todd Smith(1), Jeffrey Rosenfeld(2), Christopher Mason(3). (1) Geospiza Inc. Seattle, WA 98119, USA (2) Sackler Institute for Comparative Genomics, American Museum of Natural History, New York, NY 10024, USA (3) Weill Cornell Medical College, New York, NY 10021, USA

Kidd JM, Sampas N, Antonacci F, Graves T, Fulton R, Hayden HS, Alkan C, Malig M, Ventura M, Giannuzzi G, Kallicki J, Anderson P, Tsalenko A, Yamada NA, Tsang P, Kaul R, Wilson RK, Bruhn L, & Eichler EE (2010). Characterization of missing human genome sequences and copy-number polymorphic insertions. Nature methods, 7 (5), 365-71 PMID: 20440878

You can obtain abstracts for all of the presentations at the SFAF website.

Friday, May 20, 2011

21st Century Medicine: A Question of Ps

Last Sunday and Monday (5/15, 5/16/11) the Institute for Systems Biology (ISB) held their annual symposium. This year was the 10th annual and focused on "Systems Biology and P4 Medicine."

For those new to P4 medicine, the Ps stand for Personalized, Predictive, Preventative, and Participatory. P4 medicine is about changing our current disease oriented, reactive, approaches to those that prevent disease by increasing the predictive power of diagnostics. Because we are all different, future diagnostics need to be tailored to each individual, which also means individuals need to be more aware of their health and proactively participate in their health care. The vision of P4 medicine is that it will not only dramatically improve the quality of health care, it will significantly decrease health care costs. Hence, some folks add additional Ps to include payment and policy.

P4 medicine is an ambitious goal. In Lee Hood's closing notes he noted four significant challenges that need to be overcome to make P4 medicine a practical reality:

  1. IT challenges. In addition to working on how to transform datasets containing billions of measurements into actionable information, we need to integrate high dimensional data a wide variety of measurement systems. Reduced data will need to be presented in medical records that can be easily accessed and understood by health care providers and participants.
  2. Education. Students, scientists, doctors, individuals, and policy makers need to learn and develop an understanding of how the networks of interacting proteins and biochemicals that make us healthy or sick are regulated by our genomes and respond to environmental factors. 
  3. Big vs Small Science. Funding agencies are concerned with how to best support the research needed to create the kinds of technologies and approaches that will unlock biology's complexity to develop future diagnostics and efficacious therapies. Large-scale projects conducted over the past 10 years have made it clear that biology is extremely complex. Deciphering this complexity requires that we integrate production-orientated data collection approaches, that develop a data infrastructure, with focused research projects, run by domain experts, that explore specific ideas.  The challenge is balancing big and small science to achieve high impact goals. 
  4. Families. Understanding the genetic basis of health and disease requires the research be conducted on samples derived from families rather than randomized populations. Many families are needed to develop critical insights. However, in the U.S. our IRB (Institutional Review Boards) are considered a hinderance to enrolling individuals. 
Through the day and half conference numerous presentations explored different aspects of the above challenges. Walter Jessen at Biomarker Commons has created excellent summaries of the first and second day's presentations. For those who like raw data, the #ISB2011P4 hashtag can be used to get the symposium's tweets.

Friday, May 6, 2011

Big News! PerkinElmer Acquires Geospiza

Yesterday, May 5th, PerkinElmer announced that they have acquired Geospiza. This is exciting news and a great opportunity for Geospiza's current customers, future clients, and the company itself.

From numerous tweets, to more formal news coverage, the response has been great. Xconomy, Genome Web, Genetic Engineering & Biotechnology News (GEN), and many others have covered the story and summarized the key points.

What does it mean?

For our customers, we will continue to provide great lab (LIMS) and analysis software with excellent support. Geospiza will continue to operate in Seattle, because Seattle has a vibrant biotech and software technology environment, which has always been a great benefit for the company.

As PerkinElmer is a global organization committed to delivering advanced technology solutions to improve our health and environment, Geospiza will be able to do more interesting things and grow in exciting ways. For that, stay tuned ...

Thursday, April 28, 2011

Product Updates: GeneSifter Lab and GeneSifter Analysis Editions

Spring is here and so are new releases of the GeneSifter products. GeneSifter Lab Edition (GSLE) has been bumped up to 3.17 and GeneSifter Analysis Edition (GSAE) is now at 3.7.

What's New?

GSLE - This release includes big features along and a host of improvements. For starters, we added comprehensive inventory tracking. Now, when you configure forms to track your laboratory processes, you can add and track the use of inventory items.

Inventory items are those reagents, kits, tubes, and other bits that are used to prepare samples for analysis. GSLE makes it easy to add these items and their details like barcodes, lot numbers, vendor data, and expiration dates. Items contain arbitrary units so you can track weights as easily as volumes.

When inventory items are used in the laboratory, they can be included in steps. Each time the item is used, the amount to use can be preconfigured and GSLE will do the math for you. When the inventory item's amount drops below a threshold, GSLE can send an email that includes a link for reordering.

In addition to inventory items we increased support for the PacBio RS, and have made Sample Sheet template design completely user configurable. Sample Sheets are those files that contain the samples' names and other information needed for a data collection run. While GSLE always had good sample sheet support, vendor's frequently change formatting and needed data requirements. In some cases a new software release could be required.

The new sample sheet configuration interface eliminates the above problem, and makes it easy for labs to adapt their sheets to changes. A simple web form is used to define formatting rules and the data that will be added. GSLE tags are used to specify data fields and a search interface provides access to all fields in the database. At run time, the sample sheet is filled with the appropriate data.

View the product sheet to see the interfaces and other features.

GSAE - The new release continues to advance GSAE's data analysis capabilities. Specific features include paired-end data analysis for RNA-Seq and DNA re-sequencing applications. We've also improved the ways in which large datasets can be searched, filtered, and queried. Additional improvements include new dashboards to simplify data access and setting up analysis pipelines.

Finally, for those participating in our partner program, we continue to increase the integration between GSLE and GSAE with single sign on and data transfer features.

More details can be found in the product sheet.