Presentations
Investigating the Mechanisms of TAD Formation
Chromosome 3D structure is a blueprint that may drive epigenome maintenance. Combining with NGS, Hi-C can provide high-resolution chromatin spatial contact information on a genome-wide scale. In this talk, we will discuss recent advances in Hi-C map analysis and 3D genome folding, including:
- HiCmapTools: A new tool for efficiently accessing Hi-C maps.
- Deep learning for TAD recognition: A deep learning approach was developed to recognize TADs in Hi-C maps accurately.
- Identification of insulator elements in Drosophila: Hi-C was used to identify hundreds of insulator elements in Drosophila cells lacking the key insulator proteins CTCF and Cp190.
- TADs as 3D structural units: TADs are fundamental 3D genome units that engage in dynamic higher-order inter-TAD connections.
References
- 2023 Science Advances Topological screen identifies hundreds of Cp190-and CTCF-dependent Drosophila chromatin insulator elements
- 2022 BMC Bioinformatics HiCmapTools: a tool to access HiC contact maps
- 2021 BMC Bioinformatics Pattern recognition of topologically associating domains using deep learning
Computational protein function prediction
Biological data has grown explosively with the advance of next-generation sequencing. However, annotating protein function with wet lab experiments is time-consuming. Fortunately, computational function prediction can help wet labs formulate biological hypotheses and prioritize experiments. We have developed, GODoc, a novel and effective strategy to incorporate a training procedure into the k-nearest neighbor algorithm (instance-based learning) which is capable of solving the Gene Ontology (GO) multiple-label prediction problem, which is especially notable given the thousands of GO terms. In the CAFA3 competition (68 teams), GODoc ranks 10th in Cellular Component Ontology. In the term-centric task, GODoc performs third and is tied for first for the biofilm formation of Pseudomonas aeruginosa and the long-term memory of Drosophila melanogaster, respectively. Besides GO prediction, we present PSLCNN, a model using deep neural networks to predict protein subcellular localization for eukaryotes and prokaryotes. Compared with the state-of-the-art tools, PSLCNN achieves the best performance for prokaryotes and is comparable for eukaryotes.
References
- GODoc
- 2019 Genome biology The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens — team: NCCUCS
- 2019 BMC bioinformatics GODoc: A High-Throughput Protein Function Prediction using the Novel k-nearest-neighbor and Voting algorithms
- PSLCNN
- 2019 TAAI PSLCNN: Protein Subcellular Localization Prediction for Eukaryotes and Prokaryotes Using Deep Learning
- 2013 PLoS one Efficient and interpretable prediction of protein functional classes by Correspondence Analysis and Compact Set Relations
- 2008 Proteins PSLDoc: Protein subcellular localization prediction based on gapped-dipeptides and probabilistic latent semantic analysis
From sequence alignment to phylogenetic reconstruction: how to enrich signal to noise ratio?
Most evolutionary analyses or structure modeling are based upon pre-estimated multiple sequence alignment (MSA) models. From a computational point of view, it is too complex to estimate a correct alignment. Hence, increasing or identifying signal inside sequence alignment has intensified over the last few years. During the presentation, I would like to share our couple works on this topic.
The first part, transmembrane proteins (TMPs) constitute about 20~30% of all protein coding genes. We show how homology extension can be adapted and combined with a consistency based approach in order to significantly improve the multiple sequence alignment of alpha-helical TMPs. Then, homology and evolutionary modeling are the most common applications of MSAs. Both are known to be sensitive to the underlying MSA accuracy. We show how this problem can be partly overcome using the transitive consistency score (TCS), an extended version of the T-Coffee scoring scheme. Using this local evaluation function, we show that one can identify the most reliable portions of an MSA, as judged from BAliBASE and PREFAB structure-based reference alignments. Finally, We demonstrate that incorporating MSA induced uncertainty into bootstrap sampling can significantly increase correlation between clade correctness and its corresponding bootstrap value. Our procedure involves concatenating several alternative multiple sequence alignments of the same sequences, produced using different commonly used aligners. We then draw bootstrap replicates while favoring columns of the more unique aligner among the concatenated aligners.
References*
- PSI/TM-Coffee
- TCS
- 2015 Nucleic acids research TCS: a web server for multiple sequence alignment evaluation and phylogenetic reconstruction
- 2014 Molecular biology and evolution TCS: a new multiple sequence alignment reliability measure to estimate alignment accuracy and improve phylogenetic tree reconstruction
- wpSBOOST
* with great help and supervision by Paolo Di Tommaso and Cedric Notredame
Talks
2017
- 12.NOV 電腦不只選土豆,還會看病給藥與猜功能? — 資料科學者年會. Taipei, Taiwan
- 24.JUN Influence of alignment uncertainty on homology and phylogenetic modeling — The 25th South Taiwan Statistics Conference, National Taipei University, Taiwan
- 20.MAY High-throughput Protein Functional Prediction by Data Science Approach — The 34th Workshop on Combinatorial Mathematics and Computation Theory, Feng Chia University, Taiwan
- 20.MAY Influence of alignment uncertainty on homology and phylogenetic modeling — The 34th Workshop on Combinatorial Mathematics and Computation Theory, Feng Chia University, Taiwan
- 30.MAR Influence of alignment uncertainty on homology and phylogenetic modeling — Seminar, Taipei Medical School. Taipei, Taiwan
- 24.MAR 資訊怎麼幫助生物科技? 讓我們從人類基因體計畫聊 — Seminar, 大同高中. Taipei, Taiwan
- 02.MAR Influence of alignment uncertainty on homology and phylogenetic modeling — Seminar, TIGP. IIS, Taipei, Taiwan
- 21.FEB Influence of alignment uncertainty on homology and phylogenetic modeling — Seminar, 政治大學神經科學所. Taipei, Taiwan
2016
- 25.NOV How to enrich signal-to-noise ratio in sequence alignment? — Oral at International Symposium on Evolutionary Genomics and Bioinformatics, Kaohsiung, Taiwan
- 08.NOV How to enrich signal-to-noise ratio in sequence alignment? Homology and Sampling approaches — Seminar at 逢甲大學. Taichung, Taiwan
- 07.NOV How to enrich signal-to-noise ratio in sequence alignment? Homology and Sampling approaches — Seminar at 政治大學應用物理所. Taipei, Taiwan
- 22.SEP 如何豐富序列比對中的信噪比?利用同源性擴展與取樣方法 — Seminar at 政治大學資科系碩班 Seminar. Taichung, Taiwan
- 21.SEP How to enrich signal-to-noise ratio in sequence alignment? Sampling approaches — Seminar at 台北市立大學. Taipei, Taiwan
- 04.MAY 生物+資訊? 以染色體3級結構為例 — Seminar at 政大理學院. Taipei, Taiwan
- 24.MAR From Alignments to Phylogeny — Lecture at TIGP, Academia Sinica. Taipei, Taiwan
2015
- 19.JAN TCS: A New Multiple Sequence Alignment Reliability Measure to Estimate Alignment Accuracy and Improve Phylogenetic Tree Reconstruction — AYRCOB, National Chiao Tung University. Hsinchu, Taiwan
2013
- 03.OCT Efficient and interpretable prediction of protein functional classes by correspondence analysis and compact set relations — Seminar at PRBB Computational Genomics Seminars. Barcelona, Spain
- 01.APR Influence of alignment uncertainty on homology and phylogenetic modeling + prediction of protein functional classes — Seminar at Institute of Statistics, National Chiao Tung University. HsinChu, Taiwan