מרק לסט

אקדמי בכיר

Using machine learning methods and linguistic features in single-document extractive summarization

Alexander Dlikman, Mark Last

Extractive summarization of text documents usually consists of ranking the document sentences and extracting the top-ranked sentences subject to the summary length constraints. In this paper, we explore the contribution of various supervised learning algorithms to the sentence ranking task. For this purpose, we introduce a novel sentence ranking methodology based on the similarity score between a candidate sentence and benchmark summaries. Our experiments are performed on three benchmark summarization corpora: DUC-2002, DUC- 2007 and MultiLing-2013. The popular linear regression model achieved the best results in all evaluated datasets. Additionally, the linear regression model, which included POS (Part-of-Speech)-based features, outperformed the one with statistical features only.

שפת פרסום אנגלית
דפים 1-8
כתב עת CEUR Workshop Proceedings
כרך 1646
סטטוס פרסום פורסם - 01.01.2016

Keywords

Part-of-speech tagging
Regression
Sentence ranking
Supervised learning
Text summarization

ASJC Scopus subject areas

General Computer Science
קבצים וקישורים אחרים
Link to publication in Scopus