Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data

Nat Commun. 2020 Sep 2;11(1):4301. doi: 10.1038/s41467-020-17967-y.

Abstract

Copy-number aberrations (CNAs) and whole-genome duplications (WGDs) are frequent somatic mutations in cancer but their quantification from DNA sequencing of bulk tumor samples is challenging. Standard methods for CNA inference analyze tumor samples individually; however, DNA sequencing of multiple samples from a cancer patient has recently become more common. We introduce HATCHet (Holistic Allele-specific Tumor Copy-number Heterogeneity), an algorithm that infers allele- and clone-specific CNAs and WGDs jointly across multiple tumor samples from the same patient. We show that HATCHet outperforms current state-of-the-art methods on multi-sample DNA sequencing data that we simulate using MASCoTE (Multiple Allele-specific Simulation of Copy-number Tumor Evolution). Applying HATCHet to 84 tumor samples from 14 prostate and pancreas cancer patients, we identify subclonal CNAs and WGDs that are more plausible than previously published analyses and more consistent with somatic single-nucleotide variants (SNVs) and small indels in the same samples.

Publication types

  • Research Support, N.I.H., Extramural
  • Research Support, U.S. Gov't, Non-P.H.S.

MeSH terms

  • Breast Neoplasms / genetics*
  • Breast Neoplasms / pathology
  • DNA Copy Number Variations*
  • Datasets as Topic
  • Exome Sequencing
  • Female
  • Gene Duplication*
  • High-Throughput Nucleotide Sequencing
  • Humans
  • INDEL Mutation
  • Male
  • Mutation Rate
  • Neoplasm Metastasis / genetics
  • Pancreatic Neoplasms / genetics*
  • Pancreatic Neoplasms / pathology
  • Polymorphism, Single Nucleotide
  • Prostatic Neoplasms / genetics*
  • Prostatic Neoplasms / pathology
  • Single-Cell Analysis