• Title/Summary/Keyword: Text Visualization

Search Result 214, Processing Time 0.021 seconds

Topographic Non-negative Matrix Factorization for Topic Visualization from Text Documents (Topographic non-negative matrix factorization에 기반한 텍스트 문서로부터의 토픽 가시화)

  • Chang, Jeong-Ho;Eom, Jae-Hong;Zhang, Byoung-Tak
    • Proceedings of the Korean Information Science Society Conference
    • /
    • 2006.10b
    • /
    • pp.324-329
    • /
    • 2006
  • Non-negative matrix factorization(NMF) 기법은 음이 아닌 값으로 구성된 데이터를 두 종류의 양의 행렬의 곱의 형식으로 분할하는 데이터 분석기법으로서, 텍스트마이닝, 바이오인포매틱스, 멀티미디어 데이터 분석 등에 활용되었다. 본 연구에서는 기본 NMF 기법에 기반하여 텍스트 문서로부터 토픽을 추출하고 동시에 이를 가시적으로 도시하기 위한 Topographic NMF (TNMF) 기법을 제안한다. TNMF에 의한 토픽 가시화는 데이터를 전체적인 관점에서 보다 직관적으로 파악하는데 도움이 될 수 있다. TNMF는 생성모델 관점에서 볼 때, 2개의 은닉층을 갖는 계층적 모델로 표현할 수 있으며, 상위 은닉층에서 하위 은닉층으로의 연결은 토픽공간상에서 토픽간의 전이확률 또는 이웃함수를 정의한다. TNMF에서의 학습은 전이확률값의 연속적 스케줄링 과정 속에서 반복적 파리미터 갱신 과정을 통해 학습이 이루어지는데, 파라미터 갱신은 기본 NMF 기반 학습 과정으로부터 유사한 형태로 유도될 수 있음을 보인다. 추가적으로 Probabilistic LSA에 기초한 토픽 가시화 기법 및 희소(sparse)한 해(解) 도출을 목적으로 한 non-smooth NMF 기법과의 연관성을 분석, 제시한다. NIPS 학회 논문 데이터에 대한 실험을 통해 제안된 방법론이 문서 내에 내재된 토픽들을 효과적으로 가시화 할 수 있음을 제시한다.

  • PDF

An Analysis of Indications of Meridians in DongUiBoGam Using Data Mining (데이터마이닝을 이용한 동의보감에서 경락의 주치특성 분석)

  • Chae, Younbyoung;Ryu, Yeonhee;Jung, Won-Mo
    • Korean Journal of Acupuncture
    • /
    • v.36 no.4
    • /
    • pp.292-299
    • /
    • 2019
  • Objectives : DongUiBoGam is one of the representative medical literatures in Korea. We used text mining methods and analyzed the characteristics of the indications of each meridian in the second chapter of DongUiBoGam, WaeHyeong, which addresses external body elements. We also visualized the relationships between the meridians and the disease sites. Methods : Using the term frequency-inverse document frequency (TF-IDF) method, we quantified values regarding the indications of each meridian according to the frequency of the occurrences of 14 meridians and 14 disease sites. The spatial patterns of the indications of each meridian were visualized on a human body template according to the TF-IDF values. Using hierarchical clustering methods, twelve meridians were clustered into four groups based on the TF-IDF distributions of each meridian. Results : TF-IDF values of each meridian showed different constellation patterns at different disease sites. The spatial patterns of the indications of each meridian were similar to the route of the corresponding meridian. Conclusions : The present study identified spatial patterns between meridians and disease sites. These findings suggest that the constellations of the indications of meridians are primarily associated with the lines of the meridian system. We strongly believe that these findings will further the current understanding of indications of acupoints and meridians.

Big Data Smoothing and Outlier Removal for Patent Big Data Analysis

  • Choi, JunHyeog;Jun, Sunghae
    • Journal of the Korea Society of Computer and Information
    • /
    • v.21 no.8
    • /
    • pp.77-84
    • /
    • 2016
  • In general statistical analysis, we need to make a normal assumption. If this assumption is not satisfied, we cannot expect a good result of statistical data analysis. Most of statistical methods processing the outlier and noise also need to the assumption. But the assumption is not satisfied in big data because of its large volume and heterogeneity. So we propose a methodology based on box-plot and data smoothing for controling outlier and noise in big data analysis. The proposed methodology is not dependent upon the normal assumption. In addition, we select patent documents as target domain of big data because patent big data analysis is a important issue in management of technology. We analyze patent documents using big data learning methods for technology analysis. The collected patent data from patent databases on the world are preprocessed and analyzed by text mining and statistics. But the most researches about patent big data analysis did not consider the outlier and noise problem. This problem decreases the accuracy of prediction and increases the variance of parameter estimation. In this paper, we check the existence of the outlier and noise in patent big data. To know whether the outlier is or not in the patent big data, we use box-plot and smoothing visualization. We use the patent documents related to three dimensional printing technology to illustrate how the proposed methodology can be used for finding the existence of noise in the searched patent big data.

A Novel Interactive Power Electronics Seminar (iPES) Developed at the Swiss Federal Institute of Technology (ETH) Zurich

  • Drofenik, Uwe;Kolar, Johann W.
    • Journal of Power Electronics
    • /
    • v.2 no.4
    • /
    • pp.250-257
    • /
    • 2002
  • This paper introduces the Interactive Power Electronics Seminar - iPES - a new software package for teaching of fundamentals of power electronic circuits and systems. iPES is constituted by HTML text with Java applets for interactive animation, circuit design and simulation and visualization of electromagnetic fields and thermal issues in power electronics. It does comprise an easy-to-use self-explaining graphical user interface. The software does need just a standard web-browser, i.e. no installations are required. iPES can be accessed via the World Wide Web or from a CD-ROM in a stand-alone PC by students and professionals. Due to the underlying software technology iPES is very flexible and could be used for on-line learning and could easily be integrated into an e-learning platform. The aim of this paper Is to give an introduction to the iPES-project and to show the different areas covered. The e- learning software is available at no costs at $\underline{www.ipes.ethz.ch}$ in English, German, Japanese, Korean, Chinese and Spanish. The project is still under development and the web page is updated in about 4 weeks intervals.

A New Communication Network Model for Chat Agents in Virtual Space

  • Kim, Jong-Woo;Ji, Seong-Hyun;Kim, Seon-Yeong;Cho, Hwan-Gue
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • v.5 no.2
    • /
    • pp.287-312
    • /
    • 2011
  • Internet chat programs and instant messaging services are becoming increasingly popular among Internet users. One of the crucial issues with Internet chat is how to manage the corresponding pairs of questions and answers in a sequence of conversations. Although many novel methodologies have been introduced to cope with this problem, most are poor in managing interruptions, organizing turn-taking, and conveying comprehension. The Internet environment is recently evolving into a 3D environment, but the problems with managing chat dialogues with the standard 2D text-based chat have remained. Therefore, we propose a more realistic communication model for chat agents in 3D virtual space in this paper. First, we propose a new method to measure the capacity of communication between chat agents and a novel visualization method to depict the hierarchical structure of chat dialogues. In addition, we are concerned with communication networks for virtual people (avatars) living in virtual worlds. In this paper we consider a microscopic aspect of a social network in a relatively short period of time. Our experiments show that our model is highly effective in a virtual chat environment, and the communication network based on our model greatly facilitates investigation of a very large and complicated communication network.

A Study on the Issue Lifecycle through the Analysis of News Texts - A Case of Samsung Galaxy Note 7 - (신문기사 분석을 통한 이슈 라이프사이클에 관한 연구 - 삼성 갤럭시노트7 사례 -)

  • Heo, Pil Hee;Kim, Yang Sok;Lee, Choong Kwon
    • Smart Media Journal
    • /
    • v.7 no.4
    • /
    • pp.99-105
    • /
    • 2018
  • It is often the case that products or services on the market are causing problems, which hurt the business and image of the company. Responding appropriately to the problem and minimizing the damage is very important to business organizations. This study collected and analyzed the news articles related to the recall of the Galaxy Note 7, which was developed and launched by Samsung Electronics, one of the smartphone market leaders. Based on the issue lifecycle, the characteristics of the news were expressed by stages and the contents of the news were analyzed and visualized using association rules. The results of this study are expected to help business organizations to understand the changes and trends of issues and search for counter measures.

Favorable analysis of users through the social data analysis based on sentimental analysis (소셜데이터 감성분석을 통한 사용자의 호감도 분석)

  • Lee, Min-gyu;Sohn, Hyo-jung;Seong, Baek-min;Kim, Jong-bae
    • Proceedings of the Korean Institute of Information and Commucation Sciences Conference
    • /
    • 2014.10a
    • /
    • pp.438-440
    • /
    • 2014
  • Recently it is used commercially to actively move the data from the SNS service. Therefore, we propose a method that can accurately analyze the information related to the reputation of companies and products in real time SNS environment in this paper.Identify the relationship between words by performing morphological analysis on the text data gathered by crawling the SNS scheme. In addition, it shows the visualization to analyze statistically through a established emotional dictionary morphemes are extracted from the sentence. Here, if the extracted word is not exist in sentimental dictionary. Also, we propose the algorithm that add the word to emotional dictionary automatically.

  • PDF

A Study on Space Consumption Behavior of Contemporary Consumers -Focusing on Analysis of Social Media Big Data- (현대 소비자의 공간소비행동에 관한 연구 -소셜미디어 데이터 분석을 중심으로-)

  • Ahn, Suh Young;Koh, Ae-Ran
    • Journal of the Korean Society of Clothing and Textiles
    • /
    • v.44 no.5
    • /
    • pp.1019-1035
    • /
    • 2020
  • This study examines the millennial generation, who express themselves and share information on social media after experiencing constantly changing 'hot places' (places of interest) in contemporary cities, with the goal of analyzing space consumption behaviors. Data were collected via an Instagram crawler application developed with Python 3.4 administered to 19,262 posts using the term 'hot places' from November 1 and December 15, 2019. Issues were derived from a text mining technique using Textom 2.0; in addition, semantic network analysis using Ucinet6 and the NetDraw program were also conducted. The results are as follows. First, a frequency analysis of keywords for hot places indicated words frequently found in nouns were related to food, local names, SNS and timing. Words related to positive emotions felt in experience, and words related to behavior in hot places appeared in predicate. Based on importance, communication is the most important keyword and influenced all issues. Second, the results of visualization of semantic network analysis revealed four categories in the scope of the definition of "hot place": (1) culinary exploration, (2) atmosphere of cafés, (3) happy daily life of 'me' expressed in images, (4) emotional photos.

Analysis of Infertility Keywords in the Largest Domestic Mom Cafe Bulletin Board in Korea Using Text Mining

  • Sangmin Lee
    • Journal of Internet Computing and Services
    • /
    • v.24 no.4
    • /
    • pp.137-144
    • /
    • 2023
  • The purpose of this study is to examine consumers' perceptions of domestic infertility support policies based on infertility-related keywords and the trends of their changes. To this end, Momsholic, a mom cafe which has the most active infertility-related bulletin boards on Naver, was selected as the analysis target, and 'infertility' was selected as a keyword for data search. The data was collected for three months. In addition, network analysis and visualization were performed using R for data collection and analysis, and cross-validation was attempted using the NetDraw function of 'textom 1.0' and the UCINET6 program. As a result of the analysis, the main keywords were cost, artificial insemination, in vitro fertilization, freezing, harvest, ovulation, and how much. Next, looking at the central value of the degree of connection, it was found that the degree of connection between the words cost, cost, how much, problem, public health center, and artificial insemination was high. According to the results of this study, women who visit mom cafes due to infertility in Korea are more interested in the cost. It is believed to be closely related to infertility treatment as well as in vitro fertilization and egg freezing. Therefore, by examining keywords related toinfertility, it has academic significance in that it is possible to identify major factors that end users are interested in. Furthermore, it is possible to redefine the guidelines for domestic infertility support policies by presenting infertility support policies that reflect the factors of interest of end consumers.

Analysis of the Current Status of Edutech in Korean Language Education

  • JinHee KIM;HoSung WOO
    • Fourth Industrial Review
    • /
    • v.3 no.2
    • /
    • pp.11-17
    • /
    • 2023
  • Purpose - Recently, in the field of language education, interest in edutech has increased due to difficulties in classroom teaching due to COVID-19. Accordingly, we would like to analyze research topics related to e-learning before and after COVID-19 and examine the implications for the future Korean language education field. Research design, data, and methodology - This study organized a list of papers to be analyzed by searching for e-learning terms applicable to Korean language education in RISS. The collected data was electronically documented, keywords were extracted using text mining techniques, and word frequencies were checked, and then viewed through cloud visualization. Result - It was confirmed that research on e-learning in the field of Korean language education has increased rapidly in 2021 and 2022. In particular, extensive research on online learning methods has been actively conducted due to the difficulties of face-to-face learning in the COVID-19 era. There have been many studies on teaching and learning methods, such as flipped learning, hybrid learning, blended learning, mobile learning, and smart learning. Conclusion - Since the research so far has mainly focused on online class management methods. Therefore, future research suggests that efforts should be made to develop educational contents and teaching methods using specific ICT technologies. These efforts will contribute to advancing smart education that future education aims for.