Armin Shmilovici Leib

Senior Academic

Using a VOM model for reconstructing potential coding regions in EST sequences

Armin Shmilovici, Irad Ben-Gal

This paper presents a method for annotating coding and noncoding DNA regions by using variable order Markov (VOM) models. A main advantage in using VOM models is that their order may vary for different sequences, depending on the sequences' statistics. As a result, VOM models are more flexible with respect to model parameterization and can be trained on relatively short sequences and on low-quality datasets, such as expressed sequence tags (ESTs). The paper presents a modified VOM model for detecting and correcting insertion and deletion sequencing errors that are commonly found in ESTs. In a series of experiments the proposed method is found to be robust to random errors in these sequences.

Publication language English
Pages 49-69
Journal Computational Statistics
Volume 22
Issue number 1
Publication status Published - 01.04.2007

Keywords

Coding and noncoding DNA
Context tree
Gene annotation
Sequencing error detection and correction
Variable order Markov model

ASJC Scopus subject areas

Statistics and Probability
Statistics, Probability and Uncertainty
Computational Mathematics
Access to Document
10.1007/s00180-007-0021-8
Other files and links
Link to publication in Scopus