Get 20M+ Full-Text Papers For Less Than $1.50/day. Start a 7-Day Trial for You or Your Team.

Learn More →

Building the sequence map of the human pan-genome

Building the sequence map of the human pan-genome Here we integrate the de novo assembly of an Asian and an African genome with the NCBI reference human genome, as a step toward constructing the human pan-genome. We identified ∼5 Mb of novel sequences not present in the reference genome in each of these assemblies. Most novel sequences are individual or population specific, as revealed by their comparison to all available human DNA sequence and by PCR validation using the human genome diversity cell line panel. We found novel sequences present in patterns consistent with known human migration paths. Cross-species conservation analysis of predicted genes indicated that the novel sequences contain potentially functional coding regions. We estimate that a complete human pan-genome would contain ∼19–40 Mb of novel sequence not present in the extant reference genome. The extensive amount of novel sequence contributing to the genetic variation of the pan-genome indicates the importance of using complete genome sequencing and de novo assembly. http://www.deepdyve.com/assets/images/DeepDyve-Logo-lg.png Nature Biotechnology Springer Journals

Loading next page...
 
/lp/springer-journals/building-the-sequence-map-of-the-human-pan-genome-dUL670iFGE

References (39)

Publisher
Springer Journals
Copyright
Copyright © 2010 by Nature Publishing Group
Subject
Life Sciences; Life Sciences, general; Biotechnology; Biomedicine, general; Agriculture; Biomedical Engineering/Biotechnology; Bioinformatics
ISSN
1087-0156
eISSN
1546-1696
DOI
10.1038/nbt.1596
Publisher site
See Article on Publisher Site

Abstract

Here we integrate the de novo assembly of an Asian and an African genome with the NCBI reference human genome, as a step toward constructing the human pan-genome. We identified ∼5 Mb of novel sequences not present in the reference genome in each of these assemblies. Most novel sequences are individual or population specific, as revealed by their comparison to all available human DNA sequence and by PCR validation using the human genome diversity cell line panel. We found novel sequences present in patterns consistent with known human migration paths. Cross-species conservation analysis of predicted genes indicated that the novel sequences contain potentially functional coding regions. We estimate that a complete human pan-genome would contain ∼19–40 Mb of novel sequence not present in the extant reference genome. The extensive amount of novel sequence contributing to the genetic variation of the pan-genome indicates the importance of using complete genome sequencing and de novo assembly.

Journal

Nature BiotechnologySpringer Journals

Published: Dec 7, 2009

There are no references for this article.