Application of Machine Learning to Biomarker Discovery and Outcome Prediction in Colon Cancer using Genomic Data

Publication Date

July 2019

Location

LSU Health Medical Education Building

Document Type

Abstract

Start Date

26-7-2019 9:00 AM

End Date

26-7-2019 12:00 PM

Description

Background: Despite extensive screening campaigns, colon cancer remains the second-most common cause of cancer-related death in the United States. The recent surge of next generation sequencing of cancer genomes has led to an expanded molecular classification of colon cancer and increased our understanding of the molecular taxonomy of the disease, as genetics play a major role in its pathogenesis. However, despite the remarkable progress in mapping the genomic landscape of colon cancer, challenges remain. One of the more significant challenges is the identification of patients at high risk of developing aggressive disease that should be prioritized for treatment to improve clinical outcomes. Current prognostic and risk prediction markers include lymph invasion and tumor size, but these markers lack specificity and sensitivity. Thus, there is an urgent need for the development of more robust method for stratifying colon cancer patients according to risk to improve clinical outcomes. The application of machine learning (ML) to analysis of genomics data provides new opportunities for the development of such algorithms for patient stratification to guide treatment decisions. The objective of this investigation was to develop and apply ML to the discovery of molecular markers for stratifying patients and the prediction of patient outcome in colon cancer. Methods: To address this unmet need, we used gene expression data generated using RNA-sequencing of 509 patients from The Cancer Genome Atlas (TCGA), classified as either tumor (n = 468) or normal (n = 41). We processed and screened the data before normalization to account for differing library sizes and gene lengths. We performed supervised analyses comparing gene expressions levels in tumors and controls to identify genes associated with the disease. Using clinical information, we then sorted the patients according to the disease status at follow-up as either tumor presenting (TP, n = 77) or tumor free (TF, n = 266). We performed supervised analysis comparing gene expression levels of genes significant associated with colon cancer between the two patient groups. Significantly differentially expressed genes between TF and TP were used in four algorithms to classify patients based on disease status. The algorithms included: Naïve Bayes (NB), Very Fast Decision Tree (VFDT), Support Vector Machine (SVM), and Logistic Regression (LR). Because of the unbalanced design of the project, we applied both class-balancing and boosting meta-algorithms to all four methods to improve performance of each classifier. Accuracy was chosen as the primary evaluation metric. Results: Comparison of gene expression levels between tumors and controls revealed 13,108 (p<0.05) differentially expressed genes; only these genes were considered in the following analyses regarding disease status. Comparison of gene expression levels with respect to disease status produced 537 significantly differentially expressed genes. Application of ML to multiple subsets of these genes was executed in WEKA 3.8.2. Among the four methods used, VFDT performed the best, achieving an accuracy of 82% on 98 genes (LFC > 0.75). Conclusion: Despite limitations, including time to follow-up and disease heterogeneity, our results validate the use of genomic data to stratify patients based on risk of developing aggressive cancer, in addition to discovering clinically relevant biomarkers. Further studies are recommended to integrate transcriptome data with somatic and epigenomics data to develop more robust classifiers for potential clinical use.

Comments

Mentor: Chindo Hicks, PhD, Director of Bioinformatics Program, Department of Genetics

This document is currently not available here.

Share

COinS
 
Jul 26th, 9:00 AM Jul 26th, 12:00 PM

Application of Machine Learning to Biomarker Discovery and Outcome Prediction in Colon Cancer using Genomic Data

LSU Health Medical Education Building

Background: Despite extensive screening campaigns, colon cancer remains the second-most common cause of cancer-related death in the United States. The recent surge of next generation sequencing of cancer genomes has led to an expanded molecular classification of colon cancer and increased our understanding of the molecular taxonomy of the disease, as genetics play a major role in its pathogenesis. However, despite the remarkable progress in mapping the genomic landscape of colon cancer, challenges remain. One of the more significant challenges is the identification of patients at high risk of developing aggressive disease that should be prioritized for treatment to improve clinical outcomes. Current prognostic and risk prediction markers include lymph invasion and tumor size, but these markers lack specificity and sensitivity. Thus, there is an urgent need for the development of more robust method for stratifying colon cancer patients according to risk to improve clinical outcomes. The application of machine learning (ML) to analysis of genomics data provides new opportunities for the development of such algorithms for patient stratification to guide treatment decisions. The objective of this investigation was to develop and apply ML to the discovery of molecular markers for stratifying patients and the prediction of patient outcome in colon cancer. Methods: To address this unmet need, we used gene expression data generated using RNA-sequencing of 509 patients from The Cancer Genome Atlas (TCGA), classified as either tumor (n = 468) or normal (n = 41). We processed and screened the data before normalization to account for differing library sizes and gene lengths. We performed supervised analyses comparing gene expressions levels in tumors and controls to identify genes associated with the disease. Using clinical information, we then sorted the patients according to the disease status at follow-up as either tumor presenting (TP, n = 77) or tumor free (TF, n = 266). We performed supervised analysis comparing gene expression levels of genes significant associated with colon cancer between the two patient groups. Significantly differentially expressed genes between TF and TP were used in four algorithms to classify patients based on disease status. The algorithms included: Naïve Bayes (NB), Very Fast Decision Tree (VFDT), Support Vector Machine (SVM), and Logistic Regression (LR). Because of the unbalanced design of the project, we applied both class-balancing and boosting meta-algorithms to all four methods to improve performance of each classifier. Accuracy was chosen as the primary evaluation metric. Results: Comparison of gene expression levels between tumors and controls revealed 13,108 (p<0.05) differentially expressed genes; only these genes were considered in the following analyses regarding disease status. Comparison of gene expression levels with respect to disease status produced 537 significantly differentially expressed genes. Application of ML to multiple subsets of these genes was executed in WEKA 3.8.2. Among the four methods used, VFDT performed the best, achieving an accuracy of 82% on 98 genes (LFC > 0.75). Conclusion: Despite limitations, including time to follow-up and disease heterogeneity, our results validate the use of genomic data to stratify patients based on risk of developing aggressive cancer, in addition to discovering clinically relevant biomarkers. Further studies are recommended to integrate transcriptome data with somatic and epigenomics data to develop more robust classifiers for potential clinical use.