Origination and evolution of orphan genes and de novo genes in the genome of Caenorhabditis elegans
- 70 Downloads
Orphan genes that lack detectable homologues in other lineages could contribute to a variety of biological functions. However, their origination and function mechanisms remain largely unknown. Herein, through a comprehensive and systematic computational pipeline, we identified 893 orphan genes in the lineage of C. elegans, of which only a low fraction (0.9%) were derived from transposon elements. Six new protein-coding genes that de novo originated from non-coding DNA sequences in the genome of C. elegans were also identified. The authenticity and functionality of these orphan genes and de novo genes are supported by three lines of evidences, consisting of transcriptional data, and in silico proteomic data, and the fixation status data in wild populations. Orphan genes and de novo genes exhibited simple gene structures, such as, short in protein length, of fewer exons, and are frequently X-linked. RNA-seq data analysis showed these orphan genes are enriched with expression in embryo development and gonad, and their potential function in early development was further supported by gene ontology enrichment analysis results. Meanwhile, de novo genes were found to be with significant expression in gonad, and functional enrichment analysis of the co-expression genes of these de novo genes suggested they may be functionally involved in signaling transduction pathway and metabolism process. Our results presented the first systematic evidence on the evolution of orphan genes and de novo origin of genes in nematodes and their impacts on the functional and phenotypic evolution, and thus could shed new light on our appreciation of the importance of these new genes.
KeywordsCaenorhabditis elegans orphan genes de novo genes
We are grateful to Li Zhang for providing helpful suggestions on de novo gene identification analysis. Computing was supported by EEgrid cluster of the University of Chicago. This work was supported by National Natural Science Foundation of China (31600670 to W. Zhang, 31670851 to B. Shen).
- Arnold, A., Rahman, M.M., Lee, M.C., Muehlhaeusser, S., Katic, I., Hess, D., Scheckel, C., Wright, J.E., Stetak, A., Boag, P.R., et al. (2014). Functional characterization of C. elegans Y-box-binding proteins reveals tissue-specific functions and a critical role in the formation of polysomes. Nucleic Acids Res 42, 13353–13369.CrossRefGoogle Scholar
- Babraham Institute. (2013). FastQC: A quality control tool for high throughput sequence data. Babraham Bioinforma.Google Scholar
- Boutet, E., Lieberherr, D., Tognolli, M., Schneider, M., and Bairoch A. (2007). UniProtKB/Swiss-Prot. Methods Mol Biol 406, 89–112.Google Scholar
- Katju, V., and Lynch, M.. (2003). The structure and early evolution of recently arisen gene duplicates in the Caenorhabditis elegans genome. Genetics 165, 1793–1803.Google Scholar
- Krueger F. (2016). Trim Galore. Babraham Bioinforma.Google Scholar
- Rödelsperger, C., Streit, A., and Sommer, R.J. (2013). Structure, function and evolution of the nematode genome. In eLS (Chichester, UK: John Wiley & Sons, Ltd).Google Scholar
- Susumu O. (1970). Evolution by Gene Duplication (Springer).Google Scholar
- Williams, S. (1996). Pearson’s correlation coefficient. N Z Med J 109, 38.Google Scholar
- Zhang, Y.E., Vibranovski, M.D., Landback, P.,. Marais, G.A.B, and Long, M. (2010b). Chromosomal redistribution of male-biased genes in mammalian evolution with two bursts of gene gain on the X chromosome. PLoS Biol 8.Google Scholar
- Zhang, W., Landback, P., Gschwend, A.R., Shen, B., and Long, M. (2015). New genes drive the evolution of gene interaction networks in the human and mouse genomes. Genome Biol 16.Google Scholar