A reliable identifier of ncRNAs based on supervised self-organizing maps with rejection

Motivation : Non-coding RNAs (ncRNAs) play important roles in many biological processes and are involved in many diseases. Their identification is an important task, and many tools exist in the literature for this purpose. However, almost all of them are focused on the discrimination of coding and ncRNAs without giving more biological insight. In this paper, we propose a new reliable method called IRSOM, based on a supervised Self-Organizing Map (SOM) with a rejection option, that overcomes these limitations. The rejection option in IRSOM improves the accuracy of the method and also allows identifing the ambiguous transcripts. Furthermore, with the visualization of the SOM, we analyze the rejected predictions and highlight the ambiguity of the transcripts.

Results : IRSOM was tested on datasets of several species from different reigns, and shown better results compared to state-of-art. The accuracy of IRSOM is always greater than 0.95 for all the species with an average specificity of 0.98 and an average sensitivity of 0.99. Besides, IRSOM is fast (it takes around 254 seconds to analyze a dataset of 147 000 transcripts) and is able to handle very large datasets.



The code repository of IRSOM is available following the link below :


IRSOM sources


Datasets used in the paper are available following the link below :


IRSOM datasets

How to cite IRSOM:
  • Platon L., Zehraoui F., Bendahmane A., Tahi F. IRSOM, a reliable identifier of ncRNAs based on supervised self-organizing maps with rejection. Bioinformatics, Volume 34, Issue 17, 01 September 2018, Pages i620–i628. doi:10.1093/bioinformatics/bty572
For any questions, comments or suggestions about IRSOM, please feel free to contact: fariza.tahi@ibisc.univ-evry.fr