מרק לסט

אקדמי בכיר

Sentence Compression as a Supervised Learning with a Rich Feature Space

Elena Churkin, Mark Last, Marina Litvak, Natalia Vanetik

We present a novel supervised approach to sentence compression, based on classification and removal of word sequences generated from subtrees of the original sentence dependency tree. Our system may use any known classifier like Support Vector Machines or Logistic Model Tree to identify word sequences that can be removed without compromising the grammatical correctness of the compressed sentence. We trained our system using several classifiers on a small annotated dataset of 100 sentences, which included around 1500 manually labeled subtrees (removal candidates) represented by 25 features. The highest cross-validation classification accuracy of 80% was obtained with the SMO (Normalized Poly Kernel) algorithm. We evaluated the readability and the informativeness of the sentences compressed by the SMO-based classification model with the help of human raters using a separate benchmark dataset of 200 sentences.

שפת פרסום אנגלית
דפים 261-271
סטטוס פרסום פורסם - 01.01.2023

Keywords

Sentence compression
Supervised learning
Syntactic dependencies

ASJC Scopus subject areas

Theoretical Computer Science
General Computer Science
גישה למסמך
10.1007/978-3-031-23804-8_21
קבצים וקישורים אחרים
Link to publication in Scopus