
מרק לסט
Using machine learning methods and linguistic features in single-document extractive summarization
Extractive summarization of text documents usually consists of ranking the document sentences and extracting the top-ranked sentences subject to the summary length constraints. In this paper, we explore the contribution of various supervised learning algorithms to the sentence ranking task. For this purpose, we introduce a novel sentence ranking methodology based on the similarity score between a candidate sentence and benchmark summaries. Our experiments are performed on three benchmark summarization corpora: DUC-2002, DUC- 2007 and MultiLing-2013. The popular linear regression model achieved the best results in all evaluated datasets. Additionally, the linear regression model, which included POS (Part-of-Speech)-based features, outperformed the one with statistical features only.
| שפת פרסום | אנגלית |
| דפים | 1-8 |
| כתב עת | CEUR Workshop Proceedings |
| כרך | 1646 |
| סטטוס פרסום | פורסם - 01.01.2016 |