• Title/Summary/Keyword: text mining technique

Search Result 222, Processing Time 0.02 seconds

Analysis of the National Police Agency business trends using text mining (텍스트 마이닝 기법을 이용한 경찰청 업무 트렌드 분석)

  • Sun, Hyunseok;Lim, Changwon
    • The Korean Journal of Applied Statistics
    • /
    • v.32 no.2
    • /
    • pp.301-317
    • /
    • 2019
  • There has been significant research conducted on how to discover various insights through text data using statistical techniques. In this study we analyzed text data produced by the Korean National Police Agency to identify trends in the work by year and compare work characteristics among local authorities by identifying distinctive keywords in documents produced by each local authority. A preprocessing according to the characteristics of each data was conducted and the frequency of words for each document was calculated in order to draw a meaningful conclusion. The simple term frequency shown in the document is difficult to describe the characteristics of the keywords; therefore, the frequency for each term was newly calculated using the term frequency-inverse document frequency weights. The L2 norm normalization technique was used to compare the frequency of words. The analysis can be used as basic data that can be newly for future police work improvement policies and as a method to improve the efficiency of the police service that also help identify a demand for improvements in indoor work.

A Study on Environmental research Trends by Information and Communications Technologies using Text-mining Technology (텍스트 마이닝 기법을 이용한 환경 분야의 ICT 활용 연구 동향 분석)

  • Park, Boyoung;Oh, Kwan-Young;Lee, Jung-Ho;Yoon, Jung-Ho;Lee, Seung Kuk;Lee, Moung-Jin
    • Korean Journal of Remote Sensing
    • /
    • v.33 no.2
    • /
    • pp.189-199
    • /
    • 2017
  • Thisstudy quantitatively analyzed the research trendsin the use ofICT ofthe environmental field using the text mining technique. To that end, the study collected 359 papers published in the past two decades(1996-2015)from the National Digital Science Library (NDSL) using 38 environment-related keywords and 16 ICT-related keywords. It processed the natural languages of the environment and ICT fields in the papers and reorganized the classification system into the unit of corpus. It conducted the text mining analysis techniques of frequency analysis, keyword analysis and the association rule analysis of keywords, based on the above-mentioned keywords of the classification system. As a result, the frequency of the keywords of 'general environment' and 'climate' accounted for 77 % of the total proportion and the keywords of 'public convergence service' and 'industrial convergence service' in the ICT field took up approximately 30 % of the total proportion. According to the time series analysis, the researches using ICT in the environmental field rapidly increased over the past 5 years (2011-2015) and the number of such researches more than doubled compared to the past (1996-2010). Based on the environmental field with generated association rules among the keywords, it was identified that the keyword 'general environment' was using 16 ICT-based technologies and 'climate' was using 14 ICT-based technologies.

A Trend Analysis of Agricultural and Food Marketing Studies Using Text-mining Technique (텍스트마이닝 기법을 이용한 국내 농식품유통 연구동향 분석)

  • Yoo, Li-Na;Hwang, Su-Chul
    • Journal of the Korea Academia-Industrial cooperation Society
    • /
    • v.18 no.10
    • /
    • pp.215-226
    • /
    • 2017
  • This study analyzed trends in agricultural and food marketing studies from 1984 to 2015 using text-mining techniques. Text-mining is a part of Big-data analysis, which is an effective tool to objectively process large amounts of information based on categorization and trend analysis. In the present study, frequency analysis, topic analysis and association rules were conducted. Titles of agricultural and food marketing studies in four journals and reports were used for placing the analysis. The results showed that 1,126 total theses related to agricultural and food marketing could be categorized into six subjects. There were significant changes in research trends before and after the 2000s. While research before 2000s focused on farm and wholesale level marketing, research after the 2000s mainly covered consumption, (processed)food, exports and imports. Local food and school meals are new subjects that are increasingly being studied. Issues regarding agricultural supply and demand were the only subjects investigated in policy research studies. Interest in agricultural supply and demand was lost after the 2000s. A number of studies after the 2010s analyzed consumption, primarily consumption trends and consumer behavior.

Occupational Therapy in Long-Term Care Insurance For the Elderly Using Text Mining (텍스트 마이닝을 활용한 노인장기요양보험에서의 작업치료: 2007-2018년)

  • Cho, Min Seok;Baek, Soon Hyung;Park, Eom-Ji;Park, Soo Hee
    • Journal of Society of Occupational Therapy for the Aged and Dementia
    • /
    • v.12 no.2
    • /
    • pp.67-74
    • /
    • 2018
  • Objective : The purpose of this study is to quantitatively analyze the role of occupational therapy in long - term care insurance for the elderly using text mining, one of the big data analysis techniques. Method : For the analysis of newspaper articles, "Long - Term Care Insurance for the Elderly + Occupational Therapy for the Elderly" was collected after the period from 2007 to 208. Naver, which has a high share of the domestic search engine, utilized the database of Naver News by utilizing Textom, a web crawling tool. After collecting the article title and original text of 510 news data from the collection of the elderly long term care insurance + occupational therapy search, we analyzed the article frequency and key words by year. Result : In terms of the frequency of articles published by year, the number of articles published in 2015 and 2017 was the highest with 70 articles (13.7%), and the top 10 terms of the key word analysis showed the highest frequency of 'dementia' (344) In terms of key words, dementia, treatment, hospital, health, service, rehabilitation, facilities, institution, grade, elderly, professional, salary, industrial complex and people are related. Conclusion : In this study, it is meaningful that the textual mining technique was used to more objectively confirm the social needs and the role of the occupational therapist for the dementia and rehabilitation in the related key keywords based on the media reporting trend of the elderly long - term care insurance for 11 years. Based on the results of this study, future research should expand research field and period and supplement the research methodology through various analysis methods according to the year.

Changes in mathematics pedagogical lexicons: Extension research of the International Classroom Lexicon using a text mining approach (수학 교수학적 어휘의 변화: 텍스트 마이닝 기법을 이용한 교실수업 어휘 연구의 확장)

  • Lee, Gima;Kim, Hee-jeong
    • The Mathematical Education
    • /
    • v.61 no.4
    • /
    • pp.559-579
    • /
    • 2022
  • Research on lexicon and language provides insights into the interests, values and practices of a community where individuals use the language. The International Classroom Lexicon Project, in which ten countries participated, identified own country's mathematics teaching and learning lexicons by investigating mathematics classroom instruction from teachers' perspectives in a speaking-oriented community. This study, as an extension of the International Classroom Lexicon Project research, investigated pedagogical lexicons used in 「Mathematics and Education」 journals specialized for Korean professional mathematics teachers published by the Korean Society of Teachers of Mathematics. Using the text mining approach, we also traced how these pedegogical lexicons have changed quantitatively over the past 10 years with a diachronic perspective. As a results, several novel terms were found in the writing-oriented community, which were not identified in the speaking-oriented community. In addition, we could discover some pedagogical lexicons have increased statistically significantly and some lexicons appeared(increased) rapidly across years. This implies the teacher community's values and zeitgeist by reflecting these changes in the sociocultural, incidental and social changing (i.e., periodical change) contexts. This study has value as a first step in understanding zeitgeist for mathematics education in Korean mathematics teacher community according to changes of times over the past 10 years. Also, this study contributes to the methodological insights: the text mining technique provides a methodological contribution to researching changes in interests, values and zeitgeist according to these changes in the times.

The Effects of Cultural Factors in Tourists' Restaurant Satisfaction: Using Text Mining and Online Reviews (문화적 요인이 관광객의 음식점 만족도에 미치는 영향: 텍스트 마이닝과 온라인 리뷰를 활용하여)

  • Jiajia Meng;Gee-Woo Bock;Han-Min Kim
    • Information Systems Review
    • /
    • v.25 no.1
    • /
    • pp.145-164
    • /
    • 2023
  • The proliferation of online reviews on dining experiences has significantly affected consumers' choices of restaurants, especially overseas. Food quality, service, ambiance, and price have been identified as specific attributes for the choice of a restaurant in prior studies. In addition to these four representative attributes, cultural factors, which may also significantly impact the choice of a restaurant for tourists, in particular, have not received much attention in previous studies. This study employs the text mining technique to analyze over 10,000 online reviews of 76 Korean restaurants posted by Chinese tourists on dianping.com to explore the influence of cultural factors on the consumer's choice of restaurants in the overseas travel context. The findings reveal that "Hallyu (Korean Wave)" influences Chinese tourists' dining experiences in Korea and their satisfaction. Moreover, Korean food-related words, such as cold noodle, bibimbap, rice cake, pig trotters, and kimchi stew, appeared across all the review topics. Our findings contribute to the existing tourism and hospitality literature by identifying the critical role of cultural factors on consumers', especially tourists', satisfaction with the choice of a restaurant using text mining. The findings also provide practical guidance to restaurant owners in Korea to attract more Chinese tourists.

An Exploratory Study of e-Learning Satisfaction: A Mixed Methods of Text Mining and Interview Approaches (이러닝 만족도 증진을 위한 탐색적 연구: 텍스트 마이닝과 인터뷰 혼합방법론)

  • Sun-Gyu Lee;Soobin Choi;Hee-Woong Kim
    • Information Systems Review
    • /
    • v.21 no.1
    • /
    • pp.39-59
    • /
    • 2019
  • E-learning has improved the educational effect by making it possible to learn anytime and anywhere by escaping the traditional infusion education. As the use of e-learning system increases with the increasing popularity of e-learning, it has become important to measure e-learning satisfaction. In this study, we used the mixed research method to identify satisfaction factors of e-learning. The mixed research method is to perform both qualitative research and quantitative research at the same time. As a quantitative research, we collected reviews in Udemy.com by text mining. Then we classified high and low rated lectures and applied topic modeling technique to derive factors from reviews. Also, this study conducted an in-depth 1:1 interview on e-learning learners as a qualitative research. By combining these results, we were able to derive factors of e-learning satisfaction and dissatisfaction. Based on these factors, we suggested ways to improve e-learning satisfaction. In contrast to the fact that survey-based research was mainly conducted in the past, this study collects actual data by text mining. The academic significance of this study is that the results of the topic modeling are combined with the factor based on the information system success model.

Trends identification of species distribution modeling study in Korea using text-mining technique (텍스트마이닝을 활용한 종분포모형의 국내 연구 동향 파악)

  • Dong-Joo Kim;Yong Sung Kwon;Na-Yeon Han;Do-Hun Lee
    • Korean Journal of Environmental Biology
    • /
    • v.41 no.4
    • /
    • pp.413-426
    • /
    • 2023
  • Species distribution model (SDM) is used to preserve biodiversity and climate change impact. To evaluate biodiversity, various studies are being conducted to utilize and apply SDM. However, there is insufficient research to provide useful information by identifying the current status and recent trends of SDM research and discussing implications for future research. This study analyzed the trends and flow of academic papers, in the use of SDM, published in academic journals in South Korea and provides basic information that can be used for related research in the future. The current state and trends of SDM research were presented using philological methods and text-mining. The papers on SDM have been published 148 times between 1998 and 2023 with 115 (77.7%) papers published since 2015. MaxEnt model was the most widely used, and plant was the main target species. Most of the publications were related to species distribution and evaluation, and climate change. In text mining, the term 'Climate change' emerged as the most frequent keyword and most studies seem to consider biodiversity changes caused by climate change as a topic. In the future, the use of SDM requires several considerations such as selecting the models that are most suitable for various conditions, ensemble models, development of quantitative input variables, and improving the collection system of field survey data. Promoting these methods could help SDM serve as valuable scientific tools for addressing national policy issues like biodiversity conservation and climate change.

Towards Effective Entity Extraction of Scientific Documents using Discriminative Linguistic Features

  • Hwang, Sangwon;Hong, Jang-Eui;Nam, Young-Kwang
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • v.13 no.3
    • /
    • pp.1639-1658
    • /
    • 2019
  • Named entity recognition (NER) is an important technique for improving the performance of data mining and big data analytics. In previous studies, NER systems have been employed to identify named-entities using statistical methods based on prior information or linguistic features; however, such methods are limited in that they are unable to recognize unregistered or unlearned objects. In this paper, a method is proposed to extract objects, such as technologies, theories, or person names, by analyzing the collocation relationship between certain words that simultaneously appear around specific words in the abstracts of academic journals. The method is executed as follows. First, the data is preprocessed using data cleaning and sentence detection to separate the text into single sentences. Then, part-of-speech (POS) tagging is applied to the individual sentences. After this, the appearance and collocation information of the other POS tags is analyzed, excluding the entity candidates, such as nouns. Finally, an entity recognition model is created based on analyzing and classifying the information in the sentences.

Generating and Controlling an Interlinking Network of Technical Terms to Enhance Data Utilization (데이터 활용률 제고를 위한 기술 용어의 상호 네트워크 생성과 통제)

  • Jeong, Do-Heon
    • Journal of the Korean Society for information Management
    • /
    • v.35 no.1
    • /
    • pp.157-182
    • /
    • 2018
  • As data management and processing techniques have been developed rapidly in the era of big data, nowadays a lot of business companies and researchers have been interested in long tail data which were ignored in the past. This study proposes methods for generating and controlling a network of technical terms based on text mining technique to enhance data utilization in the distribution of long tail theory. Especially, an edit distance technique of text mining has given us efficient methods to automatically create an interlinking network of technical terms in the scholarly field. We have also used linked open data system to gather experimental data to improve data utilization and proposed effective methods to use data of LOD systems and algorithm to recognize patterns of terms. Finally, the performance evaluation test of the network of technical terms has shown that the proposed methods were useful to enhance the rate of data utilization.