• Title/Summary/Keyword: LDA word distribution

Search Result 21, Processing Time 0.022 seconds

Topic Extraction and Classification Method Based on Comment Sets

  • Tan, Xiaodong
    • Journal of Information Processing Systems
    • /
    • v.16 no.2
    • /
    • pp.329-342
    • /
    • 2020
  • In recent years, emotional text classification is one of the essential research contents in the field of natural language processing. It has been widely used in the sentiment analysis of commodities like hotels, and other commentary corpus. This paper proposes an improved W-LDA (weighted latent Dirichlet allocation) topic model to improve the shortcomings of traditional LDA topic models. In the process of the topic of word sampling and its word distribution expectation calculation of the Gibbs of the W-LDA topic model. An average weighted value is adopted to avoid topic-related words from being submerged by high-frequency words, to improve the distinction of the topic. It further integrates the highest classification of the algorithm of support vector machine based on the extracted high-quality document-topic distribution and topic-word vectors. Finally, an efficient integration method is constructed for the analysis and extraction of emotional words, topic distribution calculations, and sentiment classification. Through tests on real teaching evaluation data and test set of public comment set, the results show that the method proposed in the paper has distinct advantages compared with other two typical algorithms in terms of subject differentiation, classification precision, and F1-measure.

Feature Expansion based on LDA Word Distribution for Performance Improvement of Informal Document Classification (비격식 문서 분류 성능 개선을 위한 LDA 단어 분포 기반의 자질 확장)

  • Lee, Hokyung;Yang, Seon;Ko, Youngjoong
    • Journal of KIISE
    • /
    • v.43 no.9
    • /
    • pp.1008-1014
    • /
    • 2016
  • Data such as Twitter, Facebook, and customer reviews belong to the informal document group, whereas, newspapers that have grammar correction step belong to the formal document group. Finding consistent rules or patterns in informal documents is difficult, as compared to formal documents. Hence, there is a need for additional approaches to improve informal document analysis. In this study, we classified Twitter data, a representative informal document, into ten categories. To improve performance, we revised and expanded features based on LDA(Latent Dirichlet allocation) word distribution. Using LDA top-ranked words, the other words were separated or bundled, and the feature set was thus expanded repeatedly. Finally, we conducted document classification with the expanded features. Experimental results indicated that the proposed method improved the micro-averaged F1-score of 7.11%p, as compared to the results before the feature expansion step.

Topic Modeling of News Article Related to Franchise Regulation Using LDA (LDA 를 이용한 '프랜차이즈 규제' 관련 뉴스기사 토픽모델링)

  • YANG, Woo-Ryeong;YANG, Hoe Chang
    • The Korean Journal of Franchise Management
    • /
    • v.13 no.4
    • /
    • pp.1-12
    • /
    • 2022
  • Purpose: In 2020, the franchise industry accomplished a significant growth compared to the previous year, as the number of franchise companies increased by 9.0% while the number of franchise brands increased by 12.5%. Despite growth in size, the Korean franchise industry underwent many negative incidents, such as franchise ownership sales to private equity funds, that led to deterioration of businesses. From this point of view, this study aims to make various proposals to help policy makers develop franchise industry policies by analyzing trends of the current and previous presidential administrations' franchise policies and regulations using newspaper articles. Research design, data and methodology: A total of 7,439 articles registered in Naver API from February 25, 2013 to November 29, 2021 were extracted. Among them, 34 unrelated video articles were deleted, and a total of 7,405 articles from both administrations were used for analysis. The R package was used for word frequency analysis, word clouding, word correlation analysis, and LDA (Latent Dirichlet Allocation) topic modeling. Results: The keyword frequency analysis shows that the most frequently mentioned keywords during the previous administration include 'no-brand', 'major company', 'bill', 'business field', and 'SMEs', and those mentioned during the current administration include 'industry' and 'policy'. As a result of LDA topic modeling, 9 topics such as 'global startups' and 'job creation' from the previous administration, and 10 topics such as 'franchise business' and 'distribution industry' from the current administration were derived. The results of LDAvis showed that the previous administration operated a policy based on mutual growth of large and small businesses rather than hostile regulations in the franchise business, whereas the current administration extended the regulation related to franchise business to the employment sector. Conclusions: The analysis of past two administrations' franchise policy, it can be suggested that franchisors and franchisees may complement each other in developing the Fair Transactions in Franchise Business Act and achieving balanced growth. Moreover, political support is needed for sound development of franchisors. Limitations and future research suggestions are presented at the end of this study.

A Method on Associated Document Recommendation with Word Correlation Weights (단어 연관성 가중치를 적용한 연관 문서 추천 방법)

  • Kim, Seonmi;Na, InSeop;Shin, Juhyun
    • Journal of Korea Multimedia Society
    • /
    • v.22 no.2
    • /
    • pp.250-259
    • /
    • 2019
  • Big data processing technology and artificial intelligence (AI) are increasingly attracting attention. Natural language processing is an important research area of artificial intelligence. In this paper, we use Korean news articles to extract topic distributions in documents and word distribution vectors in topics through LDA-based Topic Modeling. Then, we use Word2vec to vector words, and generate a weight matrix to derive the relevance SCORE considering the semantic relationship between the words. We propose a way to recommend documents in order of high score.

Overseas Research Trends Related to 'Research Ethics' Using LDA Topic Modeling

  • YANG, Woo-Ryeong;YANG, Hoe-Chang
    • Journal of Research and Publication Ethics
    • /
    • v.3 no.1
    • /
    • pp.7-11
    • /
    • 2022
  • Purpose: The purpose of this study is to derive clues about the development direction of research ethics and areas of interest which has recently become a social issue in Korea by confirming overseas research trends. Research design, data and methodology: We collected 2,760 articles in scienceON, which including 'research ethics' in their paper. For analysis, frequency analysis, word clouding, keyword association analysis, and LDA topic modeling were used. Results: It was confirmed that many of the papers were published in medical, bio, pharmaceutical, and nursing journals and its interest has been continuously increasing. From word frequency analysis, many words of medical fields such as health, clinical, and patient was confirmed. From topic modeling, 7 topics were extracted such as ethical policy development and human clinical ethics. Conclusions: We founded that overseas research trends on research ethics are related to basic aspects than Korea. This means that a fundamental approach to ethics and the application of strict standards can become the basis for cultivating an overall ethical awareness. Therefore, academic discussions on the application of strict standards for publishing ethics and conducting researches in various fields where community awareness and social consensus are necessary for overall ethical awareness.

Exploring Depression Research Trends Using BERTopic and LDA

  • Woo-Ryeong, YANG;Hoe-Chang, YANG
    • The Korean Journal of Food & Health Convergence
    • /
    • v.9 no.1
    • /
    • pp.19-28
    • /
    • 2023
  • The purpose of this study is to explore which areas have been more interested in depression research in Korea through analysis of academic papers related to depression, and then to provide insights that can solve future depression problems. 1,032 papers searched with the keyword "depression" in scienceON were analyzed using Python 3.7 for word frequency analysis, word co-occurrence analysis, BERTopic, LDA, and OLS regression analysis. The results of word frequency and co-occurrence frequency analysis showed that related words were composed around words such as patient, disorder and symptom. As a result of topic modeling, a total of 13 topics including 'childhood depression' and 'eating anxiety' were derived. And it has been identified as a topic of interest that 'suicidal thoughts', 'treatment', 'occupational health', and 'health treatment program' were statistically significant topics, while 'child depression' and 'female treatment' were relatively less. As a result of the analysis of research trends, future research will not only study physiological and psychological factors but also social and environmental causes, as well as it was suggested that various collaborative studies of experts in academia were needed such as convergence and complex perspectives for depression relief and treatment.

Research Trend Analysis on Customer Satisfaction in Service Field Using BERTopic and LDA

  • YANG, Woo-Ryeong;YANG, Hoe-Chang
    • The Journal of Economics, Marketing and Management
    • /
    • v.10 no.6
    • /
    • pp.27-37
    • /
    • 2022
  • Purpose: The purpose of this study is to derive various ways to realize customer satisfaction for the development of the service industry by exploring research trends related to customer satisfaction, which is presented as an important goal in the service industry. Research design, data and methodology: To this end, 1,456 papers with English abstracts using scienceON were used for analysis. Using Python 3.7, word frequency and co-occurrence analysis were confirmed, and topics related to research trends were classified through BERTopic and LDA. Results: As a result of word frequency and co-occurrence frequency analysis, words such as quality, intention, and loyalty appeared frequently. As a result of BERTopic and LDA, 11 topics such as 'catering service' and 'brand justice' were derived. As a result of trend analysis, it was confirmed that 'brand justice' and 'internet shopping' are emerging as relatively important research topics, but CRM is less interested. Conclusions: The results of this study showed that the 7P marketing strategy is working to some extent. Therefore, it is proposed to conduct research related to acquisition of good customers through service price, customer lifetime value application, and customer segmentation that are expected to be needed for the development of the service industry.

A Study on Leadership Trends from the Perspective of Domestic Researcher's Using BERTopic and LDA

  • Sung-Su, SHIN;Hoe-Chang, Yang
    • East Asian Journal of Business Economics (EAJBE)
    • /
    • v.11 no.1
    • /
    • pp.53-71
    • /
    • 2023
  • Purpose - This study aims to find clues necessary for the direction of leadership development suitable for the current situation by exploring the direction in which leadership has been studied from the perspective of domestic researchers, along with the arrangement of leadership theories studied in various ways. Research design, data, and methodology - A total of 7,425 papers were obtained due to the search, and 5,810 papers with English abstracts were used for analysis. For analysis, word frequency analysis, word clouding, and co-occurrence were confirmed using Python 3.7. In addition, after classifying topics related to research trends through BERTopic and LDA, trends were identified through dynamic topic modeling and OLS regression analysis. Result - As a result of the BERTopic, 14 topics such as 'Leadership management and performance' and 'Sports leadership' were derived. As a result of conducting LDA on 1,976 outliers, five topics were derived. As a result of trend analysis on topics by year, it was confirmed that five topics, such as 'military police leadership' received relative attention. Conclusion - Through the results of this study, a study on the reinterpretation of past leadership studies, a study on LMX with an expanded perspective, and a study on integrated leadership sub-factors of modern leadership theory were proposed.

Research Trend Analysis of the Retail Industry: Focusing on the Department Store (유통업태 연구동향 분석: 백화점을 중심으로)

  • Hoe-Chang YANG
    • The Journal of Economics, Marketing and Management
    • /
    • v.11 no.5
    • /
    • pp.45-55
    • /
    • 2023
  • Purpose: As one of the continuous studies on the offline distribution industry, the purpose of this study is to find ways for offline stores to respond to the growth of online shopping by identifying research trends on department stores. Research design, data and methodology: To this end, this study conducted word frequency analysis, word co-occurrence frequency analysis, BERTopic, LDA, and dynamic topic modeling using Python 3.7 on a total of 551 English abstracts searched with the keyword 'department store' in scienceON as of October 10, 2022. Results: The results of word frequency analysis and co-occurrence frequency analysis revealed that research related to department stores frequently focuses on factors such as customers, consumers, products, satisfaction, services, and quality. BERTopic and LDA analyses identified five topics, including 'store image,' with 'shopping information' showing relatively high interest, while 'sales systems' were observed to have relatively lower interest. Conclusions: Based on the results of this study, it was concluded that research related to department stores has so far been conducted in a limited scope, and it is insufficient to provide clues for department stores to secure competitiveness against online platforms. Therefore, it is suggested that additional research be conducted on topics such as the true role of department stores in the retail industry, consumer reinterpretation, customer value and lifetime value, department stores as future retail spaces, ethical management, and transparent ESG management.

Online Shopping Research Trend Analysis Using BERTopic and LDA

  • Yoon-Hwang, JU;Woo-Ryeong, YANG;Hoe-Chang, YANG
    • The Journal of Economics, Marketing and Management
    • /
    • v.11 no.1
    • /
    • pp.21-30
    • /
    • 2023
  • Purpose: As one of the ongoing studies on the distribution industry, the purpose of this study is to identify the research trends on online shopping so far to propose not only the development of online shopping companies but also the possibility of coexistence between online and offline retailers and the development of the distribution industry. Research design, data and methodology: In this study, the English abstracts of 645 papers on online shopping registered in scienceON were obtained. For the analysis through BERTopic and LDA using Python 3.7 and identifying which topics were interesting to researchers. Results: As a result of word frequency analysis and co-occurrence analysis, it was found that studies related to online shopping were frequently conducted on factors such as products, services, and shopping malls. As a result of BERTopic, five topics such as 'service quality' and 'sales strategy' were derived, and as a result of LDA, three topics including 'purchase experience' were derived. It was confirmed that 'Customer Recommendation' and 'Fashion Mall' showed relatively high interest, and 'Sales Strategy' showed relatively low interest. Conclusions: It was suggested that more diverse studies related to the online shopping mall platform, sales content, and usage influencing factors are needed to develop the online shopping industry.