• 제목/요약/키워드: Text analysis

검색결과 3,326건 처리시간 0.029초

기술 문헌 분석 테스트베드 툴킷 개발 (Developing a Test-Bed Toolkit for Scientific Document Analysis)

  • 최성필;송사광;정한민
    • 한국콘텐츠학회논문지
    • /
    • 제12권8호
    • /
    • pp.13-19
    • /
    • 2012
  • 본 논문은 논문, 특허, 연구보고서 등과 같은 다양한 과학 기술 문헌에 포함된 기술 지식을 효과적으로 추출하는데 필요한 텍스트 분석 엔진들의 효과적인 모니터링 및 성능 최적화를 위한 테스트베드 도구를 소개한다. 이 도구는 과학 기술 분야의 전문 용어를 비롯한 인명, 지명, 기관명 등을 자동으로 인식하는 기술 개체 인식 엔진을 위한 테스트베드와 인식된 기술 개체 간의 의미적 연관 관계를 자동으로 추출하는 기술개체 간 관계 추출 테스트베드로 구성되어 있다. 이를 활용함으로써 사용자 및 개발자들은 기술 문헌 분석 엔진의 실행 모니터링은 물론 오류 분석을 효율적으로 수행할 수 있다.

섬유소재 분야 특허 기술 동향 분석: DETM & STM 텍스트마이닝 방법론 활용 (Research of Patent Technology Trends in Textile Materials: Text Mining Methodology Using DETM & STM)

  • 이현상;조보근;오세환;하성호
    • 한국정보시스템학회지:정보시스템연구
    • /
    • 제30권3호
    • /
    • pp.201-216
    • /
    • 2021
  • Purpose The purpose of this study is to analyze the trend of patent technology in textile materials using text mining methodology based on Dynamic Embedded Topic Model and Structural Topic Model. It is expected that this study will have positive impact on revitalizing and developing textile materials industry as finding out technology trends. Design/methodology/approach The data used in this study is 866 domestic patent text data in textile material from 1974 to 2020. In order to analyze technology trends from various aspect, Dynamic Embedded Topic Model and Structural Topic Model mechanism were used. The word embedding technique used in DETM is the GloVe technique. For Stable learning of topic modeling, amortized variational inference was performed based on the Recurrent Neural Network. Findings As a result of this analysis, it was found that 'manufacture' topics had the largest share among the six topics. Keyword trend analysis found the fact that natural and nanotechnology have recently been attracting attention. The metadata analysis results showed that manufacture technologies could have a high probability of patent registration in entire time series, but the analysis results in recent years showed that the trend of elasticity and safety technology is increasing.

The Impact of Transforming Unstructured Data into Structured Data on a Churn Prediction Model for Loan Customers

  • Jung, Hoon;Lee, Bong Gyou
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • 제14권12호
    • /
    • pp.4706-4724
    • /
    • 2020
  • With various structured data, such as the company size, loan balance, and savings accounts, the voice of customer (VOC), which is text data containing contact history and counseling details was analyzed in this study. To analyze unstructured data, the term frequency-inverse document frequency (TF-IDF) analysis, semantic network analysis, sentiment analysis, and a convolutional neural network (CNN) were implemented. A performance comparison of the models revealed that the predictive model using the CNN provided the best performance with regard to predictive power, followed by the model using the TF-IDF, and then the model using semantic network analysis. In particular, a character-level CNN and a word-level CNN were developed separately, and the character-level CNN exhibited better performance, according to an analysis for the Korean language. Moreover, a systematic selection model for optimal text mining techniques was proposed, suggesting which analytical technique is appropriate for analyzing text data depending on the context. This study also provides evidence that the results of previous studies, indicating that individual customers leave when their loyalty and switching cost are low, are also applicable to corporate customers and suggests that VOC data indicating customers' needs are very effective for predicting their behavior.

팬데믹 시기의 패션 테크놀로지에 관한 시각 - 텍스트 마이닝과 내용 분석을 중심으로 - (Perspectives on Fashion Technology during the Pandemic Era - A Mixed Methods Approach Using Text Mining and Content Analysis -)

  • 김미경;임은혁
    • 한국의류산업학회지
    • /
    • 제24권5호
    • /
    • pp.545-556
    • /
    • 2022
  • To overcome the pandemic, a new strategy for innovation is in demand throughout the value chains of the fashion industry that emphasize the importance of fashion technology. Accordingly, as various viewpoints and fields of debate are unfolding to consider the direction of change led by fashion technology, it is necessary to make an active value judgment precedent by understanding the differences between various opinions. This study aims to derive keywords from fashion technology used during the pandemic, to infer the characteristics of each type of perspective and to understand their characteristics. For the research, this study combines text mining analysis and content analysis. Text mining analysis is used to find statistical patterns by collecting keywords from big data from online media, and content analysis is used to interpret the data qualitatively. After analyzing the results of this study, the following observations are made. First, the perspective of positive acceptance seeks to maximize the perception and sensory action of fashion through technology; this amplifies experience, an opportunity for innovation and efficiency. Second, critical vigilance highlights the side effects of radical changes in fashion technology, characterized by concerns about capital-centered polarization, threats to human rights, and infringement of creative thinking. Lastly, the perspective of gradual adoption is the gradual convergence of technologies, characterized by the pursuit of an appropriate balance.

TextRank 알고리즘을 이용한 음악 가사 요약 기법 (Music Lyrics Summarization Method using TextRank Algorithm)

  • 손지영;신용태
    • 한국멀티미디어학회논문지
    • /
    • 제21권1호
    • /
    • pp.45-50
    • /
    • 2018
  • This research paper describes how to summarize music lyrics using the TextRank algorithm. This method can summarize music lyrics as important lyrics. Therefore, we recommend music more effectively than analyzing the number of words and recommending music.

Study of Analyzing Outcome of Building and Introducing System for Preserving Full-Text of e-Journal

  • Kim, Kwang-Young;Kim, Soon-Young;Kim, Hwan-Min
    • International Journal of Knowledge Content Development & Technology
    • /
    • 제2권2호
    • /
    • pp.5-16
    • /
    • 2012
  • Today, most researchers conduct their studies through the full-text of e-journals. Therefore, an important base for domestic development of science and technology is to obtain the full-text of quality e-journals by overseas researchers and to provide it to Korea's researchers. This study aims to build a system based on the National Archiving Center for the full-text of e-journals and to make a service system for providing them to the public by acquiring the full-text of quality overseas e-journals. To do this, an analysis was made of the outcome of introducing such a system for full-text of e-journals in comparison with the investment. As a result, 112 more institutions, that is, from 47 institutions to 159 institutions, have introduced the system as of 2012, and the number of downloaded full-texts increased at least 2.17 times.

한 손을 이용한 스마트폰 터치키 문자입력에서 선호손의 수행도 분석 (Performance Analysis of Text Entry with Preferred One Hand using Smart Phone Touch-keyboard)

  • 류태범
    • 대한인간공학회지
    • /
    • 제30권1호
    • /
    • pp.259-264
    • /
    • 2011
  • Does preferred hand show better performance than non-preferred hand in smart phone text entry using one hand. Is the performance of subjects who use left-preferred hand in smart phone text entry worse than that of others who use right preferred hand among the right handed. This study tried to address these two questions. Thirty young male undergraduate students typed a text using a smart phone which has a touch-based QWERTY keyboard two times with both hands, right and left hand, respectively. The completion time, errors were measured in the text entry tasks. All of participants were right handed, but half of them preferred right hand if they have to use one hand in smart phone text entry and other half preferred left hand. The percentage that preferred hand has better performance than non-preferred hand in smart phone text entry using one hand is less than 90% for right-preferred hand and less than 70% for left-preferred hand. The performance of left hand preferred students is not worse than that of the right hand preferred in one hand text entry of smart phone.

Arabic Text Clustering Methods and Suggested Solutions for Theme-Based Quran Clustering: Analysis of Literature

  • Bsoul, Qusay;Abdul Salam, Rosalina;Atwan, Jaffar;Jawarneh, Malik
    • Journal of Information Science Theory and Practice
    • /
    • 제9권4호
    • /
    • pp.15-34
    • /
    • 2021
  • Text clustering is one of the most commonly used methods for detecting themes or types of documents. Text clustering is used in many fields, but its effectiveness is still not sufficient to be used for the understanding of Arabic text, especially with respect to terms extraction, unsupervised feature selection, and clustering algorithms. In most cases, terms extraction focuses on nouns. Clustering simplifies the understanding of an Arabic text like the text of the Quran; it is important not only for Muslims but for all people who want to know more about Islam. This paper discusses the complexity and limitations of Arabic text clustering in the Quran based on their themes. Unsupervised feature selection does not consider the relationships between the selected features. One weakness of clustering algorithms is that the selection of the optimal initial centroid still depends on chances and manual settings. Consequently, this paper reviews literature about the three major stages of Arabic clustering: terms extraction, unsupervised feature selection, and clustering. Six experiments were conducted to demonstrate previously un-discussed problems related to the metrics used for feature selection and clustering. Suggestions to improve clustering of the Quran based on themes are presented and discussed.

소설텍스트의 난이도 조정 방안 연구 -이효석의 「메밀꽃 필 무렵」을 중심으로- (This study revises Lee Hyo-seok's The Buckwheat Season, utilizing Novel Corpus, intermediate learners' level)

  • 황혜란
    • 한국어교육
    • /
    • 제29권4호
    • /
    • pp.255-294
    • /
    • 2018
  • The Buckwheat Season, evaluated as the best of Lee Hyo-seok's literature, is one of the short stories that represent Korean literature. However, vivid literary expressions such as lyrical and beautiful depictions, figurative expressions and dialects, which show the Korean beauty, rather make learners have difficulty and become a factor that fails in reading comprehension. Thus, it is necessary to revise and present the text modified for the learners' language level. The methods of revising a literary text include the revision of linguistic elements such as cryptic vocabulary or sentence structure and the revision of the composition of the text, e.g. suggestion of characters or plot, or insertion of illustration. The methods of revising the language of the text can be divided into methods of simplification and detailing. However, in the process of revising the text, many depend on the adapter's subjective perception, not revising it with objective criteria. This paper revised the text, utilizing by the Academy of Korean Studies, , and the by the National Institute of Korean Language to secure objectivity in revising the text.

Identification of the Minimum Legible Text Size for Group-View Display of the Main Control Room in Radioactive Waste Facility

  • Jung, Kihyo;Lee, Baekhee;Chang, Yoon;Jung, Ilho;You, Heecheon
    • 대한인간공학회지
    • /
    • 제36권3호
    • /
    • pp.213-219
    • /
    • 2017
  • Objective: The present study identified the minimum legible text size by an experiment for eight combinations of background and text colors, which will be used in designing visual information on group-view display (GVD). Background: Information on minimum legible text size is needed to design the visual information presented on GVD in a radioactive waste control room. Method: The experiment was conducted for 22 male participants (age: mean = 37, SD = 6.7; visual acuity: over 0.8) who were recruited by considering demographic characteristics of current control room operators. Eight combinations of background and text colors were considered and the minimum legible text size was determined for each combination by applying the method of limits, one of psychophysical methods. Results: The minimum legible text size was significantly different in accordance with the combination of background and text colors. Statistical analysis results showed that luminance contrast and color contrast between background and text influenced the minimum legible text sizes. Conclusion: This study concluded that the minimum legible text size is 8 minute of arc for various combinations of background and text colors. Application: The minimum legible text size identified in the present study can be utilized in designing visual information on GVD at the main control room in a radioactive waste facility.