• 제목/요약/키워드: Topic Evaluation

검색결과 406건 처리시간 0.024초

토픽 식별성 향상을 위한 키워드 재구성 기법 (Keyword Reorganization Techniques for Improving the Identifiability of Topics)

  • 윤여일;김남규
    • 한국IT서비스학회지
    • /
    • 제18권4호
    • /
    • pp.135-149
    • /
    • 2019
  • Recently, there are many researches for extracting meaningful information from large amount of text data. Among various applications to extract information from text, topic modeling which express latent topics as a group of keywords is mainly used. Topic modeling presents several topic keywords by term/topic weight and the quality of those keywords are usually evaluated through coherence which implies the similarity of those keywords. However, the topic quality evaluation method based only on the similarity of keywords has its limitations because it is difficult to describe the content of a topic accurately enough with just a set of similar words. In this research, therefore, we propose topic keywords reorganizing method to improve the identifiability of topics. To reorganize topic keywords, each document first needs to be labeled with one representative topic which can be extracted from traditional topic modeling. After that, classification rules for classifying each document into a corresponding label are generated, and new topic keywords are extracted based on the classification rules. To evaluated the performance our method, we performed an experiment on 1,000 news articles. From the experiment, we confirmed that the keywords extracted from our proposed method have better identifiability than traditional topic keywords.

Hot Topic Discovery across Social Networks Based on Improved LDA Model

  • Liu, Chang;Hu, RuiLin
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • 제15권11호
    • /
    • pp.3935-3949
    • /
    • 2021
  • With the rapid development of Internet and big data technology, various online social network platforms have been established, producing massive information every day. Hot topic discovery aims to dig out meaningful content that users commonly concern about from the massive information on the Internet. Most of the existing hot topic discovery methods focus on a single network data source, and can hardly grasp hot spots as a whole, nor meet the challenges of text sparsity and topic hotness evaluation in cross-network scenarios. This paper proposes a novel hot topic discovery method across social network based on an im-proved LDA model, which first integrates the text information from multiple social network platforms into a unified data set, then obtains the potential topic distribution in the text through the improved LDA model. Finally, it adopts a heat evaluation method based on the word frequency of topic label words to take the latent topic with the highest heat value as a hot topic. This paper obtains data from the online social networks and constructs a cross-network topic discovery data set. The experimental results demonstrate the superiority of the proposed method compared to baseline methods.

A Comparative Study on Research Strategies for the Architectural Design Evaluation

  • Han, Seung-Hoon;Moon, Jin-Woo
    • Architectural research
    • /
    • 제12권2호
    • /
    • pp.41-52
    • /
    • 2010
  • The aim of this paper is to evaluate the methodological strategies for the architectural design field of study mainly focused on qualitative and quantitative research designs. Firstly, this paper addresses the characteristics of the six approaches including their methodological aspects in general. Each strategy is assessed by the different approaches to give full insights in it. Secondly, it distinguishes the differences among six research approaches especially derived from qualitative and quantitative research designs. The differences are discussed in terms of the strengths and weaknesses of each strategy. Finally, this paper attempts to discuss about possible applications for introducing approaches to the research topic with which the exemplified research topic, a design evaluation system, deals. To investigate the applicability of the design methods employed, the following topic has been stated; how to develop a design evaluation system and what to be considered for unfolding the thrown topic in terms of strategic approaches in the field of architectural design researches reviewed through the study.

Topic Extraction and Classification Method Based on Comment Sets

  • Tan, Xiaodong
    • Journal of Information Processing Systems
    • /
    • 제16권2호
    • /
    • pp.329-342
    • /
    • 2020
  • In recent years, emotional text classification is one of the essential research contents in the field of natural language processing. It has been widely used in the sentiment analysis of commodities like hotels, and other commentary corpus. This paper proposes an improved W-LDA (weighted latent Dirichlet allocation) topic model to improve the shortcomings of traditional LDA topic models. In the process of the topic of word sampling and its word distribution expectation calculation of the Gibbs of the W-LDA topic model. An average weighted value is adopted to avoid topic-related words from being submerged by high-frequency words, to improve the distinction of the topic. It further integrates the highest classification of the algorithm of support vector machine based on the extracted high-quality document-topic distribution and topic-word vectors. Finally, an efficient integration method is constructed for the analysis and extraction of emotional words, topic distribution calculations, and sentiment classification. Through tests on real teaching evaluation data and test set of public comment set, the results show that the method proposed in the paper has distinct advantages compared with other two typical algorithms in terms of subject differentiation, classification precision, and F1-measure.

Exploratory Study of Developing a Synchronization-Based Approach for Multi-step Discovery of Knowledge Structures

  • Yu, So Young
    • Journal of Information Science Theory and Practice
    • /
    • 제2권2호
    • /
    • pp.16-32
    • /
    • 2014
  • As Topic Modeling has been applied in increasingly various domains, the difficulty in naming and characterizing topics also has been recognized more. This study, therefore, explores an approach of combining text mining with network analysis in a multi-step approach. The concept of synchronization was applied to re-assign the top author keywords in more than one topic category, in order to improve the visibility of the topic-author keyword network, and to increase the topical cohesion in each topic. The suggested approach was applied using 16,548 articles with 2,881 unique author keywords in construction and building engineering indexed by KSCI. As a result, it was revealed that the combined approach could improve both the visibility of the topic-author keyword map and topical cohesion in most of the detected topic categories. There should be more cases of applying the approach in various domains for generalization and advancement of the approach. Also, more sophisticated evaluation methods should also be necessary to develop the suggested approach.

토픽 모델링을 활용한 교양 ICT 활용과정 서술형 강의평가 분석 (Analysis of Descriptive Lecture Evaluation on Liberal Arts ICT utilization using Topic Modeling)

  • 김효숙
    • Journal of Platform Technology
    • /
    • 제8권1호
    • /
    • pp.33-40
    • /
    • 2020
  • 본 연구의 목적은 교양 ICT활용 과정의 서술형 강의 평가에 대해 텍스트 마이닝의 토픽모델링 분석을 실시하여 수강생의 강의 선택 요인과 강의에 대한 긍정적·부정적 요소 파악을 하고자 하는데 있다. 이를 위해 M 대학교의 2019년 2학기에 개설된 ICT활용 과정 강의에 대해 '강의를 신청한 이유', '강의에서 개선되어야 할 점'과 '강의에서 좋았던 점'에 대한 데이터 전처리부터 키워드 빈도 분석, 워드 클라우드 시각화 및 토픽 모델링 분석을 실시하였다. 연구결과 M 대학의 2019년 2학기 ICT활용 과정은 자격증 취득을 위해 강의를 신청하며, 동시에 자격증을 취득할 수 있어 강의가 좋았다는 긍정적 분석을 알 수 있다. 부정적 요소로 강의실 사용 환경 불편에 대한 것을 알 수 있다.

  • PDF

Generative probabilistic model with Dirichlet prior distribution for similarity analysis of research topic

  • Milyahilu, John;Kim, Jong Nam
    • 한국멀티미디어학회논문지
    • /
    • 제23권4호
    • /
    • pp.595-602
    • /
    • 2020
  • We propose a generative probabilistic model with Dirichlet prior distribution for topic modeling and text similarity analysis. It assigns a topic and calculates text correlation between documents within a corpus. It also provides posterior probabilities that are assigned to each topic of a document based on the prior distribution in the corpus. We then present a Gibbs sampling algorithm for inference about the posterior distribution and compute text correlation among 50 abstracts from the papers published by IEEE. We also conduct a supervised learning to set a benchmark that justifies the performance of the LDA (Latent Dirichlet Allocation). The experiments show that the accuracy for topic assignment to a certain document is 76% for LDA. The results for supervised learning show the accuracy of 61%, the precision of 93% and the f1-score of 96%. A discussion for experimental results indicates a thorough justification based on probabilities, distributions, evaluation metrics and correlation coefficients with respect to topic assignment.

무한 사전 온라인 LDA 토픽 모델에서 의미적 연관성을 사용한 토픽 확장 (Topic Expansion based on Infinite Vocabulary Online LDA Topic Model using Semantic Correlation Information)

  • 곽창욱;김선중;박성배;김권양
    • 정보과학회 컴퓨팅의 실제 논문지
    • /
    • 제22권9호
    • /
    • pp.461-466
    • /
    • 2016
  • 토픽 확장은 학습된 토픽의 질을 향상시키기 위해 추가적인 외부 데이터를 반영하여 점진적으로 토픽을 확장하는 방법이다. 기존의 온라인 학습 토픽 모델에서는 외부 데이터를 확장에 사용될 경우, 새로운 단어가 기존의 학습된 모델에 반영되지 않는다는 문제가 있었다. 본 논문에서는 무한 사전 온라인 LDA 토픽 모델을 이용하여 외부 데이터를 반영한 토픽 모델 확장 방법을 연구하였다. 토픽 확장 학습에서는 기존에 형성된 토픽과 추가된 외부 데이터의 단어와 유사도를 반영하여 토픽을 확장한다. 실험에서는 기존의 토픽 확장 모델들과 비교하였다. 비교 결과, 제안한 방법에서 외부 연관 문서 단어를 토픽 모델에 반영하기 때문에 대본 토픽이 다루지 못한 정보들을 토픽에 포함할 수 있었다. 또한, 일관성 평가에서도 비교 모델보다 뛰어난 성능을 나타냈다.

토픽모델링을 활용한 국내 수학과 교육과정 연구 동향 분석 : 1997년부터 2019년까지 게재된 국내 수학교육 학술지 논문을 중심으로 (An analysis of domestic research trends of mathematics curriculum research through topic modeling: Focused on domestic journals published from 1997 to 2019)

  • 손태권;이광호
    • 한국수학교육학회지시리즈A:수학교육
    • /
    • 제59권3호
    • /
    • pp.201-216
    • /
    • 2020
  • 본 연구는 1997년부터 2019년까지 KCI 등재지에 게재된 493편의 국내 수학과 교육과정 논문을 LDA 토픽 모델링을 사용하여 연구 동향을 분석하였다. 그 결과, 국내 수학과 교육과정 연구는 8개의 토픽으로 분류할 수 있었으며 그 중 '교육과정 이행과 평가'의 비중이 가장 낮았다. 또한 교육과정 적용 시기에 따라 토픽들은 다르게 출현했으며 수학과 교육과정에서 강조하는 중점 방향과 부합하는 경향성을 보였다. 이러한 결과를 바탕으로 향후 수학과 교육과정의 발전을 위한 시사점들을 도출하였다.

Customer Service Evaluation based on Online Text Analytics: Sentiment Analysis and Structural Topic Modeling

  • 박경배;하성호
    • 한국정보시스템학회지:정보시스템연구
    • /
    • 제26권4호
    • /
    • pp.327-353
    • /
    • 2017
  • Purpose Social media such as social network services, online forums, and customer reviews have produced a plethora amount of information online. Yet, the information deluge has created both opportunities and challenges at the same time. This research particularly focuses on the challenges in order to discover and track the service defects over time derived by mining publicly available online customer reviews. Design/methodology/approach Synthesizing the streams of research from text analytics, we apply two stages of methods of sentiment analysis and structural topic model incorporating meta-information buried in review texts into the topics. Findings As a result, our study reveals that the research framework effectively leverages textual information to detect, prioritize, and categorize service defects by considering the moving trend over time. Our approach also highlights several implications theoretically and practically of how methods in computational linguistics can offer enriched insights by leveraging the online medium.