Robust high-dimensional bioinformatics data streams mining by ODR-ioVFDT

Dantong Wang, Simon Fong, Raymond K. Wong, Sabah Mohammed, Jinan Fiaidhi, Kelvin K. L. Wong

Research output: Contribution to journalArticlepeer-review

16 Citations (Scopus)

Abstract

Outlier detection in bioinformatics data streaming mining has received significant attention by research communities in recent years. The problems of how to distinguish noise from an exception and deciding whether to discard it or to devise an extra decision path for accommodating it are causing dilemma. In this paper, we propose a novel algorithm called ODR with incrementally Optimized Very Fast Decision Tree (ODR-ioVFDT) for taking care of outliers in the progress of continuous data learning. By using an adaptive interquartile-range based identification method, a tolerance threshold is set. It is then used to judge if a data of exceptional value should be included for training or otherwise. This is different from the traditional outlier detection/removal approaches which are two separate steps in processing through the data. The proposed algorithm is tested using datasets of five bioinformatics scenarios and comparing the performance of our model and other ones without ODR. The results show that ODRioVFDT has better performance in classification accuracy, kappa statistics, and time consumption. The ODR-ioVFDT applied onto bioinformatics streaming data processing for detecting and quantifying the information of life phenomena, states, characters, variables and components of the organism can help to diagnose and treat disease more effectively.
Original languageEnglish
Article number43167
Number of pages12
JournalScientific Reports
Volume7
DOIs
Publication statusPublished - 2017

Open Access - Access Right Statement

© The Author(s) 2017. This work is licensed under a Creative Commons Attribution 4.0 International License. The images or other third party material in this article are included in the article’s Creative Commons license, unless indicated otherwise in the credit line; if the material is not included under the Creative Commons license, users will need to obtain permission from the license holder to reproduce the material. To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/

Keywords

  • bioinformatics
  • data mining
  • noise

Fingerprint

Dive into the research topics of 'Robust high-dimensional bioinformatics data streams mining by ODR-ioVFDT'. Together they form a unique fingerprint.

Cite this