• Title/Summary/Keyword: sentiment dictionary

Search Result 71, Processing Time 0.033 seconds

Analyzing Contextual Polarity of Unstructured Data for Measuring Subjective Well-Being (주관적 웰빙 상태 측정을 위한 비정형 데이터의 상황기반 긍부정성 분석 방법)

  • Choi, Sukjae;Song, Yeongeun;Kwon, Ohbyung
    • Journal of Intelligence and Information Systems
    • /
    • v.22 no.1
    • /
    • pp.83-105
    • /
    • 2016
  • Measuring an individual's subjective wellbeing in an accurate, unobtrusive, and cost-effective manner is a core success factor of the wellbeing support system, which is a type of medical IT service. However, measurements with a self-report questionnaire and wearable sensors are cost-intensive and obtrusive when the wellbeing support system should be running in real-time, despite being very accurate. Recently, reasoning the state of subjective wellbeing with conventional sentiment analysis and unstructured data has been proposed as an alternative to resolve the drawbacks of the self-report questionnaire and wearable sensors. However, this approach does not consider contextual polarity, which results in lower measurement accuracy. Moreover, there is no sentimental word net or ontology for the subjective wellbeing area. Hence, this paper proposes a method to extract keywords and their contextual polarity representing the subjective wellbeing state from the unstructured text in online websites in order to improve the reasoning accuracy of the sentiment analysis. The proposed method is as follows. First, a set of general sentimental words is proposed. SentiWordNet was adopted; this is the most widely used dictionary and contains about 100,000 words such as nouns, verbs, adjectives, and adverbs with polarities from -1.0 (extremely negative) to 1.0 (extremely positive). Second, corpora on subjective wellbeing (SWB corpora) were obtained by crawling online text. A survey was conducted to prepare a learning dataset that includes an individual's opinion and the level of self-report wellness, such as stress and depression. The participants were asked to respond with their feelings about online news on two topics. Next, three data sources were extracted from the SWB corpora: demographic information, psychographic information, and the structural characteristics of the text (e.g., the number of words used in the text, simple statistics on the special characters used). These were considered to adjust the level of a specific SWB. Finally, a set of reasoning rules was generated for each wellbeing factor to estimate the SWB of an individual based on the text written by the individual. The experimental results suggested that using contextual polarity for each SWB factor (e.g., stress, depression) significantly improved the estimation accuracy compared to conventional sentiment analysis methods incorporating SentiWordNet. Even though literature is available on Korean sentiment analysis, such studies only used only a limited set of sentimental words. Due to the small number of words, many sentences are overlooked and ignored when estimating the level of sentiment. However, the proposed method can identify multiple sentiment-neutral words as sentiment words in the context of a specific SWB factor. The results also suggest that a specific type of senti-word dictionary containing contextual polarity needs to be constructed along with a dictionary based on common sense such as SenticNet. These efforts will enrich and enlarge the application area of sentic computing. The study is helpful to practitioners and managers of wellness services in that a couple of characteristics of unstructured text have been identified for improving SWB. Consistent with the literature, the results showed that the gender and age affect the SWB state when the individual is exposed to an identical queue from the online text. In addition, the length of the textual response and usage pattern of special characters were found to indicate the individual's SWB. These imply that better SWB measurement should involve collecting the textual structure and the individual's demographic conditions. In the future, the proposed method should be improved by automated identification of the contextual polarity in order to enlarge the vocabulary in a cost-effective manner.

Electronic-Composit Consumer Sentiment Index(CCSI) development by Social Bigdata Analysis (소셜빅데이터를 이용한 온라인 소비자감성지수(e-CCSI) 개발)

  • Kim, Yoosin;Hong, Sung-Gwan;Kang, Hee-Joo;Jeong, Seung-Ryul
    • Journal of Internet Computing and Services
    • /
    • v.18 no.4
    • /
    • pp.121-131
    • /
    • 2017
  • With emergence of Internet, social media, and mobile service, the consumers have actively presented their opinions and sentiment, and then it is spreading out real time as well. The user-generated text data on the Internet and social media is not only the communication text among the users but also the valuable resource to be analyzed for knowing the users' intent and sentiment. In special, economic participants have strongly asked that the social big data and its' analytics supports to recognize and forecast the economic trend in future. In this regard, the governments and the businesses are trying to apply the social big data into making the social and economic solutions. Therefore, this study aims to reveal the capability of social big data analysis for the economic use. The research proposed a social big data analysis model and an online consumer sentiment index. To test the model and index, the researchers developed an economic survey ontology, defined a sentiment dictionary for sentiment analysis, conducted classification and sentiment analysis, and calculated the online consumer sentiment index. In addition, the online consumer sentiment index was compared and validated with the composite consumer survey index of the Bank of Korea.

A Comparative Study on Using SentiWordNet for English Twitter Sentiment Analysis (영어 트위터 감성 분석을 위한 SentiWordNet 활용 기법 비교)

  • Kang, In-Su
    • Journal of the Korean Institute of Intelligent Systems
    • /
    • v.23 no.4
    • /
    • pp.317-324
    • /
    • 2013
  • Twitter sentiment analysis is to classify a tweet (message) into positive and negative sentiment class. This study deals with SentiWordNet(SWN)-based twitter sentiment analysis. SWN is a sentiment dictionary in which each sense of an English word has a positive and negative sentimental strength. There has been a variety of SWN-based sentiment feature extraction methods which typically first determine the sentiment orientation (SO) of a term in a document and then decide SO of the document from such terms' SO values. For example, for SO of a term, some calculated the maximum or average of sentiment scores of its senses, and others computed the average of the difference of positive and negative sentiment scores. For SO of a document, many researchers employ the maximum or average of terms' SO values. In addition, the above procedure may be applied to the whole set (adjective, adverb, noun, and verb) of parts-of-speech or its subset. This work provides a comparative study on SWN-based sentiment feature extraction schemes with performance evaluation on a well-known twitter dataset.

Construction of Vietnamese SentiWordNet by using Vietnamese Dictionary (베트남어 사전을 사용한 베트남어 SentiWordNet 구축)

  • Vu, Xuan-Son;Park, Seong-Bae
    • Proceedings of the Korea Information Processing Society Conference
    • /
    • 2014.04a
    • /
    • pp.745-748
    • /
    • 2014
  • SentiWordNet is an important lexical resource supporting sentiment analysis in opinion mining applications. In this paper, we propose a novel approach to construct a Vietnamese SentiWordNet (VSWN). SentiWordNet is typically generated from WordNet in which each synset has numerical scores to indicate its opinion polarities. Many previous studies obtained these scores by applying a machine learning method to WordNet. However, Vietnamese WordNet is not available unfortunately by the time of this paper. Therefore, we propose a method to construct VSWN from a Vietnamese dictionary, not from WordNet. We show the effectiveness of the proposed method by generating a VSWN with 39,561 synsets automatically. The method is experimentally tested with 266 synsets with aspect of positivity and negativity. It attains a competitive result compared with English SentiWordNet that is 0.066 and 0.052 differences for positivity and negativity sets respectively.

Emotion Analysis System for Social Media using Sentiment Dictionary including newly created word (신조어 감성사전 기반의 소셜미디어 감성분석 시스템)

  • Shin, Panseop;Oh, Hanmin
    • Proceedings of the Korean Society of Computer Information Conference
    • /
    • 2019.01a
    • /
    • pp.225-226
    • /
    • 2019
  • 오피니언 마이닝은 온라인 문서의 감성을 추출하여 분석하는 기법이다. 별도의 여론조사 없이 감성을 분석 가능하므로, 최근 활발한 연구 분야이다. 그러나 소셜미디어에는 신조어 등이 많이 포함되어 있어 기존 감성분석 시스템으로는 정확한 분석이 어려울 뿐만 아니라, 복합적인 감성에 대한 분석을 내리기에 불리하다. 이에 본 연구에서는 직관적인 감성모델을 제안하고 SNS에서 주목받는 다양한 신조어를 수용한 감성단어사전을 구축한 후, 이를 적용하여 소셜미디어에 나타나는 복합적인 감성을 분석하는 감성분석시스템을 설계한다.

  • PDF

System Design for Analysis and Evaluation of E-commerce Products Using Review Sentiment Word Analysis (리뷰 감정 분석을 통한 전자상거래 상품 분석 및 평가 시스템 설계)

  • Choi, Jieun;Ryu, Hyejin;Yu, Dabeen;Kim, Nara;Kim, Yoonhee
    • KIISE Transactions on Computing Practices
    • /
    • v.22 no.5
    • /
    • pp.209-217
    • /
    • 2016
  • As smartphone usage increases, the number of consumers who refer to review data of e-commercial products using web sites and SNS is also explosively multiplying. However, reading review data using traditional websites and SNS is time consuming. Also, it is impossible for consumers to read all the reviews. Therefore, a system that collects review data of products and conducts sentiment word analysis of the review is required to provide useful information. The majority of systems that provide such information inadequately reflect the properties of the product. In this study, we described a system that provides analysis and evaluation of e-commerce products through review sentiment words as reflected properties of the product. Furthermore, the system enables consumers to access processed information about reviews quickly and in visual format.

Exploration of Constituent Factors for Corporate Reputation and Development of Index Using Online News : Sentiment Analysis and AHP Application (온라인 뉴스를 이용한 기업평판 구성요인 탐색 및 지수 개발 연구 : 감성분석과 AHP적용)

  • Lee, Byung Hyun;Choi, Il Young;Lee, Jung Jae;Kim, Jae Kyeong;Kang, Hyun Mo
    • Journal of Information Technology Services
    • /
    • v.19 no.6
    • /
    • pp.145-159
    • /
    • 2020
  • Because of the recent development of information and communication technology, companies are exposed to various media such as blogs, social media, and YouTube. In particular, exposed news affects the company's reputation. So, while positive news can improve corporate value, negative news can lead to financial losses for the company. In this study, we redefine corporate reputation as social responsibility, vision and leadership, financial performance, products and services through existing literature, and conducted an AHP survey with a total of four components to calculate the weight of each factor. As a result of the calculation, the proportion of financial performance was the highest at 0.41, and products and services, vision and leadership, and social responsibility were the lowest. In addition, in order to measure the reputation of a company, it is classified as a component that defines online news using the LDA technique. In addition, through sentiment analysis, an index for each corporate reputation factor was derived, and the reputation index was calculated by combining it with the AHP analysis result, and Spearman ranking correlation analysis was performed to secure the validity of the research results. Therefore, the significance of this study is that the definition and importance of the constituent factors can contribute to the future planning and development direction of the company, and also contribute to the derivation of the corporate reputation index. This study is significant in that a new analysis methodology that applied AHP analysis results to sentiment analysis was suggested.

Sentiment Analysis on 'Non-maritalism Childbirth' Using Naver News Comments (네이버 뉴스 댓글을 활용한 '비혼출산'에 대한 감성분석)

  • Huh, Seyoung;Kim, Cho-Won;Cheong, Anyong;Lee, Sae Bom
    • The Journal of the Korea Contents Association
    • /
    • v.22 no.1
    • /
    • pp.74-85
    • /
    • 2022
  • Along with the change in the values of marriage and the prevalence of non-marriage in Korean society, a new form of family composition called unmarried birth or non-maritalism childbirth has appeared, and social discussion in taking place in connection with the problem of a decrease in the birthrate. Using sentiment analysis and social network analysis, this research explored how the people's sentiment and perception has changed toward 'nonmarital birth.' The data used is comments on news articles from the period of November 2020 to August 2021. As a result of the study, there were a lot of positive comments during the social issue period by marriage, whereas there were many negative comments from the policy agenda to the policy making period. As a result of co-occurrence network analysis, the topic of family norm, policy, and personal aspect appeared. This study is significant in that it revealed that negative perceptions prevailed during the policy-making process after the issue of unmarried births after the issue of unmarried births, and it became a cornerstone of social discussion on unmarried births

Movie Retrieval System by Analyzing Sentimental Keyword from User's Movie Reviews (사용자 영화평의 감정어휘 분석을 통한 영화검색시스템)

  • Oh, Sung-Ho;Kang, Shin-Jae
    • Journal of the Korea Academia-Industrial cooperation Society
    • /
    • v.14 no.3
    • /
    • pp.1422-1427
    • /
    • 2013
  • This paper proposed a movie retrieval system based on sentimental keywords extracted from user's movie reviews. At first, sentimental keyword dictionary is manually constructed by applying morphological analysis to user's movie reviews, and then keyword weights in the dictionary are calculated for each movie with TF-IDF. By using these results, the proposed system classify sentimental categories of movies and rank classified movies. Without reading any movie reviews, users can retrieve movies through queries composed by sentimental keywords.

Detection of Adverse Drug Reactions Using Drug Reviews with BERT+ Algorithm (BERT+ 알고리즘 기반 약물 리뷰를 활용한 약물 이상 반응 탐지)

  • Heo, Eun Yeong;Jeong, Hyeon-jeong;Kim, Hyon Hee
    • KIPS Transactions on Software and Data Engineering
    • /
    • v.10 no.11
    • /
    • pp.465-472
    • /
    • 2021
  • In this paper, we present an approach for detection of adverse drug reactions from drug reviews to compensate limitations of the spontaneous adverse drug reactions reporting system. Considering negative reviews usually contain adverse drug reactions, sentiment analysis on drug reviews was performed and extracted negative reviews. After then, MedDRA dictionary and named entity recognition were applied to the negative reviews to detect adverse drug reactions. For the experiment, drug reviews of Celecoxib, Naproxen, and Ibuprofen from 5 drug review sites, and analyzed. Our results showed that detection of adverse drug reactions is able to compensate to limitation of under-reporting in the spontaneous adverse drugs reactions reporting system.