איסנה לובלינסקי

אקדמי בכיר

Optimizing genomic language models for promoter prediction

a comparative study of tokenization and cross-species learning

Eyal Hadad, Noia Kogman, Lina Golan, Anva Avraham, Reut Ben-Hamo, Zhi Wei, Lior Rokach,Isana Veksler-Lublinsky

Large Language Models (LLMs) are increasingly applied to genomic tasks, yet core challenges remain concerning tokenization, evaluation, and data scarcity. This study focuses on promoter classification and systematically evaluates four tokenization methods: non-overlapping 6-mer, overlapping 6-mer, Byte Pair Encoding (BPE), and WordPiece (WPC). We show that the commonly used k-mer approach, specifically the non-overlapping variant, outperforms BPE and WPC across eight organisms, challenging assumptions derived from natural language processing. To ensure robustness, we evaluated performance under two distinct negative data strategies: positive-promoter-shuffled and random-non-promoter-fragments. Using a positional SHAP framework, we demonstrate that the model learns biologically plausible positional patterns rather than exploiting artifacts from these negative data generation processes. Furthermore, evolutionary-informed transfer learning experiments and external validation on an unseen organism reveal that training on phylogenetically related species significantly improves performance, particularly in low-data regimes. These findings underscore the significant impact of tokenization and negative data design, providing practical guidance for refining genomic classifiers.

שפת פרסום אנגלית
כתב עת NAR Genomics and Bioinformatics
כרך 8
נושא מספר 1
סטטוס פרסום פורסם - 01.03.2026
מספר מאמר lqag025

ASJC Scopus subject areas

Structural Biology
Molecular Biology
Genetics
Computer Science Applications
Applied Mathematics
גישה למסמך
10.1093/nargab/lqag025
קבצים וקישורים אחרים
Link to publication in Scopus