Plant omics data center: An integrated web repository for interspecies gene expression networks with NLP-based curation

Hajime Ohyanagi, Tomoyuki Takano, Shin Terashima, Masaaki Kobayashi, Maasa Kanno, Kyoko Morimoto, Hiromi Kanegae, Yohei Sasaki, Misa Saito, Satomi Asano, Soichi Ozaki, Toru Kudo, Koji Yokoyama, Koichiro Aya, Keita Suwabe, Go Suzuki, Koh Aoki, Yasutaka Kubo, Masao Watanabe, Makoto MatsuokaKentaro Yano

Research output: Contribution to journalArticle

37 Citations (Scopus)

Abstract

Comprehensive integration of large-scale omics resources such as genomes, transcriptomes and metabolomes will provide deeper insights into broader aspects of molecular biology. For better understanding of plant biology, we aim to construct a next-generation sequencing (NGS)-derived gene expression network (GEN) repository for a broad range of plant species. So far we have incorporated information about 745 high-quality mRNA sequencing (mRNA-Seq) samples from eight plant species (Arabidopsis thaliana, Oryza sativa, Solanum lycopersicum, Sorghum bicolor, Vitis vinifera, Solanum tuberosum, Medicago truncatula and Glycine max) from the public short read archive, digitally profiled the entire set of gene expression profiles, and drawn GENs by using correspondence analysis (CA) to take advantage of gene expression similarities. In order to understand the evolutionary significance of the GENs from multiple species, they were linked according to the orthology of each node (gene) among species. In addition to other gene expression information, functional annotation of the genes will facilitate biological comprehension. Currently we are improving the given gene annotations with natural language processing (NLP) techniques and manual curation. Here we introduce the current status of our analyses and the web database, PODC (Plant Omics Data Center; http://bioinf.mind.meiji.ac.jp/podc/), now open to the public, providing GENs, functional annotations and additional comprehensive omics resources.

Original languageEnglish
Pages (from-to)e9
JournalPlant and Cell Physiology
Volume56
Issue number1
DOIs
Publication statusPublished - Jan 1 2015

Keywords

  • Correspondence analysis
  • Database
  • Gene expression network
  • Manual curation
  • Natural language processing (NLP)
  • Omics

ASJC Scopus subject areas

  • Physiology
  • Plant Science
  • Cell Biology

Fingerprint Dive into the research topics of 'Plant omics data center: An integrated web repository for interspecies gene expression networks with NLP-based curation'. Together they form a unique fingerprint.

  • Cite this

    Ohyanagi, H., Takano, T., Terashima, S., Kobayashi, M., Kanno, M., Morimoto, K., Kanegae, H., Sasaki, Y., Saito, M., Asano, S., Ozaki, S., Kudo, T., Yokoyama, K., Aya, K., Suwabe, K., Suzuki, G., Aoki, K., Kubo, Y., Watanabe, M., ... Yano, K. (2015). Plant omics data center: An integrated web repository for interspecies gene expression networks with NLP-based curation. Plant and Cell Physiology, 56(1), e9. https://doi.org/10.1093/pcp/pcu188