Search | Korea Science

Lee, Gyoung Ho;Park, Yo-Han;Lee, Kong Joo
- KIPS Transactions on Software and Data Engineering
- /
- v.9 no.8
- /
- pp.251-258
- /
- 2020
A training dataset for text summarization consists of pairs of a document and its summary. As conventional approaches to building text summarization dataset are human labor intensive, it is not easy to construct large datasets for text summarization. A collection of news articles is one of the most popular resources for text summarization because it is easily accessible, large-scale and high-quality text. From social media news services, we can collect not only headlines and subheads of news articles but also summary descriptions that human editors write about the news articles. Approximately 425,000 pairs of news articles and their summaries are collected from social media. We implemented an automatic extractive summarizer and trained it on the dataset. The performance of the summarizer is compared with unsupervised models. The summarizer achieved better results than unsupervised models in terms of ROUGE score.
https://doi.org/10.3745/KTSDE.2020.9.8.251 인용 PDF KSCI