Browse > Article
http://dx.doi.org/10.13088/jiis.2014.20.1.133

A Study on the Effect of Using Sentiment Lexicon in Opinion Classification  

Kim, Seungwoo (Graduate School of Business IT, Kookmin University)
Kim, Namgyu (Graduate School of Business IT, Kookmin University)
Publication Information
Journal of Intelligence and Information Systems / v.20, no.1, 2014 , pp. 133-148 More about this Journal
Abstract
Recently, with the advent of various information channels, the number of has continued to grow. The main cause of this phenomenon can be found in the significant increase of unstructured data, as the use of smart devices enables users to create data in the form of text, audio, images, and video. In various types of unstructured data, the user's opinion and a variety of information is clearly expressed in text data such as news, reports, papers, and various articles. Thus, active attempts have been made to create new value by analyzing these texts. The representative techniques used in text analysis are text mining and opinion mining. These share certain important characteristics; for example, they not only use text documents as input data, but also use many natural language processing techniques such as filtering and parsing. Therefore, opinion mining is usually recognized as a sub-concept of text mining, or, in many cases, the two terms are used interchangeably in the literature. Suppose that the purpose of a certain classification analysis is to predict a positive or negative opinion contained in some documents. If we focus on the classification process, the analysis can be regarded as a traditional text mining case. However, if we observe that the target of the analysis is a positive or negative opinion, the analysis can be regarded as a typical example of opinion mining. In other words, two methods (i.e., text mining and opinion mining) are available for opinion classification. Thus, in order to distinguish between the two, a precise definition of each method is needed. In this paper, we found that it is very difficult to distinguish between the two methods clearly with respect to the purpose of analysis and the type of results. We conclude that the most definitive criterion to distinguish text mining from opinion mining is whether an analysis utilizes any kind of sentiment lexicon. We first established two prediction models, one based on opinion mining and the other on text mining. Next, we compared the main processes used by the two prediction models. Finally, we compared their prediction accuracy. We then analyzed 2,000 movie reviews. The results revealed that the prediction model based on opinion mining showed higher average prediction accuracy compared to the text mining model. Moreover, in the lift chart generated by the opinion mining based model, the prediction accuracy for the documents with strong certainty was higher than that for the documents with weak certainty. Most of all, opinion mining has a meaningful advantage in that it can reduce learning time dramatically, because a sentiment lexicon generated once can be reused in a similar application domain. Additionally, the classification results can be clearly explained by using a sentiment lexicon. This study has two limitations. First, the results of the experiments cannot be generalized, mainly because the experiment is limited to a small number of movie reviews. Additionally, various parameters in the parsing and filtering steps of the text mining may have affected the accuracy of the prediction models. However, this research contributes a performance and comparison of text mining analysis and opinion mining analysis for opinion classification. In future research, a more precise evaluation of the two methods should be made through intensive experiments.
Keywords
Sentiment Lexicon; BigData Analysis; Opinion Mining; Text Mining;
Citations & Related Records
Times Cited By KSCI : 4  (Citation Analysis)
연도 인용수 순위
1 Tsur, O., D. Davidov, and A. Rappoport, "A Great Catchy Name: Semi-Supervised Recognition of Sarcastic Sentences in Online Product Reviews," Proceedings of the International AAAI Conference on Weblogs and Social Media, (2010), 162-169.
2 Turney, P. D., "Thumbs Up or Thumbs Down?: Semantic Orientation Applied to Unsupervised Classification of Reviews," Proceedings of Annual Meeting of the Association for computational Linguistics, (2002), 417-424.
3 Albright, R., Taming Text with the SVD, SAS Institute Inc., 2006.
4 Wiebe, J., R. F. Bruce, and T. P. O'Hara, "Development and Use of a Gold-Standard Data Set for Subjectivity Classifications," Proceedings of the Association for Computational Linguistics, (1999), 246-253.
5 Yu, E., J. Kim, C. Lee, and N. Kim, "Using Ontologies for Semantic Text Mining," The Journal of Information Systems, Vol. 21, No. 3(2012), 137-161.   DOI   ScienceOn
6 Yu, E., Y. Kim, N. Kim, and S. Jeong, "Predicting the Direction of the Stock Index by Using a Domain-Specific Sentiment Dictionary," Journal of Intelligence and Information Systems, Vol. 19, No. 1(2013), 95-110.   DOI
7 Asher, N., F. Benamara, and Y. Y. Mathieu, "Distilling Opinion in Discourse: A Preliminary Study," Proceedings of the International Conference on Computational Linguistics, (2008), 7-10.
8 Cho, I. and N. Kim, "Recommending Core and Connecting Keywords of Research Area Using Social Network and Data Mining Techniques," Journal of Intelligence and Information Systems, Vol. 17, No. 1(2011), 127-138.
9 Dave, K., S. Lawrence, and D. M. Pennock, "Mining the Peanut Gallery: Opinion Extraction and Semantic Classification of Product Reviews," Proceedings of International Conference on World Wide Web, (2003), 519-528.
10 Ding, X., B. Liu, and L. Zhang, "Entity Discovery and Assignment for Opinion Mining Applications," Proceedings of ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, (2009), 1125-1134.
11 Gartner Inc., 2012 Hype Cycle for Emerging Technologies, Gartner Inc., 2012.
12 Han, J. and M. Kamber, Data Mining: Concepts and Techniques, 3rd edition., Morgan Kaufmann Publishers, 2011.
13 Hazivassiloglou, V. and K. R. McKeown, "Predicting the Semantic Orientation of Adjectives," Proceedings of Annual Meeting of the Association for Computational Linguistics, (1997), 174-181.
14 Hu, M. and B. Liu, "Mining and Summarizing Customer Reviews," Proceddings of ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, (2004).
15 Kim, S. and E. Hovy, "Determining the sentiment of Opinions," Proceedings of International Conference on Computational Linguistics, No. 1367(2004).
16 Hyun, Y., H. Han, H. Choi, J. Park, K. Lee, K. Kwahk, and N. Kim, "Methodology Using Text Analysis for Packaging R&D Information Services on Pending National Issues," Journal of Information Technology Applications & Management, Vol. 20, No. 3(2013), 231-257.
17 Jindal, N. and B. Liu, "Mining Compareative Sentences and Relations," Proceeding of National Conference on Artificial Intelligence, Vol. 2(2006), 1331-1336.
18 Kamps, J., M. Marx, R. J. Mokken, and M. D. Rijke, "Using WordNet to Measure Semantic Orientation of Adjectives," Proceedings of International Conference on Language Resources and Evaluation, Vol. 4(2004), 1115-1118.
19 Liu, B., Sentiment Analysis and Opinion Mining, Morgan and Claypool Publishers, 2012.
20 McKinsey Global Institute, Big Data: The next Frontier for Innovation, Competition, and Productivity, McKinsey and Company, 2011.
21 Narayanan, R., B. Liu, and A. Choudhary, "Sentiment Analysis of Conditional Sentences," Proceeding of Conference on Empirical Methods in Natural Language Processing, Vol. 1(2009), 180-189.
22 O'Reilly Radar Team, Big Data Now: Current Perspectives from O'Reilly Radar, O'Reilly, 2011.
23 Pang, B., L. Lee, and S. Vaithyanathan, "Thumbs Up?:Sentiment Classification using Machine Learning Techniques," Proceedings of Conference on Empirical Methods in Natural Language Processing, Vol. 10(2002), 79-86.
24 Stanvrianou, A., P. Andritsos, and N. Nicoloyannis, "Overview and Semantic Issues of Text Mining", ACM SIGMOD Record, Vol. 36, No. 3(2007), 23-34.   DOI   ScienceOn