Skip to main navigation Skip to search Skip to main content

On the sample complexity of cancer pathways identification

  • F. Vandin
  • , B. J. Raphael
  • , E. Upfal
  • Brown University

Research output: Chapter in Book/Report/Conference proceedingArticle in proceedingsResearchpeer-review

Abstract

In this work we propose a framework to analyze the sample complexity of problems that arise in the study of genomic datasets. Our framework is based on tools from combinatorial analysis and statistical learning theory that have been used for the analysis of machine learning and probably approximately correct (PAC) learning. We use our framework to analyze the problem of the identification of cancer pathways through mutual exclusivity analysis of mutations from large cancer sequencing studies. We analytically derive matching upper and lower bounds on the sample complexity of the problem, showing that sample sizes much larger than currently available may be required to identify all the cancer genes in a pathway. We also provide two algorithms to find a cancer pathway from a large genomic dataset. On simulated and cancer data, we show that our algorithms can be used to identify cancer pathways from large genomic datasets.

Original languageEnglish
Title of host publicationResearch in Computational Molecular Biology : Proceedings of the 19th Annual International Conference on Research in Computational Molecular Biology
EditorsTeresa M. Przytycka
PublisherSpringer
Publication date2015
Pages326-337
ISBN (Print)978-3-319-16705-3
ISBN (Electronic)978-3-319-16706-0
DOIs
Publication statusPublished - 2015
Event19th Annual International Conference on Research in Computational Molecular Biology - Warsaw, Poland
Duration: 12. Apr 201515. Apr 2015

Conference

Conference19th Annual International Conference on Research in Computational Molecular Biology
Country/TerritoryPoland
CityWarsaw
Period12/04/201515/04/2015
SeriesLecture Notes in Computer Science
Volume9029
ISSN0302-9743

Fingerprint

Dive into the research topics of 'On the sample complexity of cancer pathways identification'. Together they form a unique fingerprint.

Cite this