• Title/Summary/Keyword: Co-word Occurrence

Search Result 104, Processing Time 0.022 seconds

Construction of Event Networks from Large News Data Using Text Mining Techniques (텍스트 마이닝 기법을 적용한 뉴스 데이터에서의 사건 네트워크 구축)

  • Lee, Minchul;Kim, Hea-Jin
    • Journal of Intelligence and Information Systems
    • /
    • v.24 no.1
    • /
    • pp.183-203
    • /
    • 2018
  • News articles are the most suitable medium for examining the events occurring at home and abroad. Especially, as the development of information and communication technology has brought various kinds of online news media, the news about the events occurring in society has increased greatly. So automatically summarizing key events from massive amounts of news data will help users to look at many of the events at a glance. In addition, if we build and provide an event network based on the relevance of events, it will be able to greatly help the reader in understanding the current events. In this study, we propose a method for extracting event networks from large news text data. To this end, we first collected Korean political and social articles from March 2016 to March 2017, and integrated the synonyms by leaving only meaningful words through preprocessing using NPMI and Word2Vec. Latent Dirichlet allocation (LDA) topic modeling was used to calculate the subject distribution by date and to find the peak of the subject distribution and to detect the event. A total of 32 topics were extracted from the topic modeling, and the point of occurrence of the event was deduced by looking at the point at which each subject distribution surged. As a result, a total of 85 events were detected, but the final 16 events were filtered and presented using the Gaussian smoothing technique. We also calculated the relevance score between events detected to construct the event network. Using the cosine coefficient between the co-occurred events, we calculated the relevance between the events and connected the events to construct the event network. Finally, we set up the event network by setting each event to each vertex and the relevance score between events to the vertices connecting the vertices. The event network constructed in our methods helped us to sort out major events in the political and social fields in Korea that occurred in the last one year in chronological order and at the same time identify which events are related to certain events. Our approach differs from existing event detection methods in that LDA topic modeling makes it possible to easily analyze large amounts of data and to identify the relevance of events that were difficult to detect in existing event detection. We applied various text mining techniques and Word2vec technique in the text preprocessing to improve the accuracy of the extraction of proper nouns and synthetic nouns, which have been difficult in analyzing existing Korean texts, can be found. In this study, the detection and network configuration techniques of the event have the following advantages in practical application. First, LDA topic modeling, which is unsupervised learning, can easily analyze subject and topic words and distribution from huge amount of data. Also, by using the date information of the collected news articles, it is possible to express the distribution by topic in a time series. Second, we can find out the connection of events in the form of present and summarized form by calculating relevance score and constructing event network by using simultaneous occurrence of topics that are difficult to grasp in existing event detection. It can be seen from the fact that the inter-event relevance-based event network proposed in this study was actually constructed in order of occurrence time. It is also possible to identify what happened as a starting point for a series of events through the event network. The limitation of this study is that the characteristics of LDA topic modeling have different results according to the initial parameters and the number of subjects, and the subject and event name of the analysis result should be given by the subjective judgment of the researcher. Also, since each topic is assumed to be exclusive and independent, it does not take into account the relevance between themes. Subsequent studies need to calculate the relevance between events that are not covered in this study or those that belong to the same subject.

A Method for Information Source Selection using Teasaurus for Distributed Information Retrieval

  • Goto, Shoji;Ozono, Tadachika;Shintani, Toramatsu
    • Proceedings of the Korea Inteligent Information System Society Conference
    • /
    • 2001.01a
    • /
    • pp.272-277
    • /
    • 2001
  • In this paper, we describe a new method for selecting information sources in a distributed environment. Recently, there has been much research on distributed information retrieval, that is information retrieval (IR) based on a multi-database model in which the existence of multiple sources is modeled explicitly. In distributed IR, a method is needed that would enable selecting appropriate sources for users\` queries. Most existing methods use statistical data such as document frequency. These methods may select inappropriate ate sources if a query contains polysemous words. In this paper, we describe an information-source selection method using two types of thesaurus. One is a thesaurus automatically constructed from documents in a source. The other is a hand-crafted general-purpose thesaurus(e.g. WordNet). Terms used in documents in a source differ from one another and the meanings of a term differ depending on th situation in which the term is used. The difference is a characteristic of the source. In our method, the meanings of a term are distinguished between by the relationship between the term and other terms, and the relationship appear in the co-occurrence-based thesaurus. In this paper, we describe an algorithm for evaluating a usefulness of a source for a query based on a thesaurus. For a practical application of our method, we have developed Papits, a multi-agent-based in formation sharing system. An experiment of selection shows that our method is effective for selecting appropriate sources.

  • PDF

An Improved Homonym Disambiguation Model based on Bayes Theory (Bayes 정리에 기반한 개선된 동형이의어 분별 모텔)

  • 김창환;이왕우
    • Journal of the Korea Computer Industry Society
    • /
    • v.2 no.12
    • /
    • pp.1581-1590
    • /
    • 2001
  • This paper asserted more developmental model of WSD(word sense disambiguation) than J. Hur(2000)'s WSD model. This model suggested an improved statistical homonym disambiguation Model based on Bayes Theory. This paper using semantic information(co-occurrence data) obtained from definitions of part of speech(POS) tagged UMRD-S(Ulsan university Machine Readable Dictionary(Semantic Tagged)). we extracted semantic features in the context as nouns, predicates and adverbs from the definitions in the korean dictionary. In this research, we make an experiment with the accuracy of WSD system about major nine homonym nouns and new seven homonym predicates supplementary. The inner experimental result showed average accuracy of 98.32% with regard to the most Nine homonym nouns and 99.53% for the Seven homonym predicates. An Addition, we save test on Korean Information Base and ETRI's POS tagged corpus. This external experimental result showed average accuracy of 84.42% with regard to the most Nine nouns over unsupervised learning sentences from Korean Information Base and ETRI Corpus, 70.81 % accuracy rate for the Seven predicates from Sejong Project phrase part tagging corpus (3.5 million phrases) too.

  • PDF

Multi-class Support Vector Machines Model Based Clustering for Hierarchical Document Categorization in Big Data Environment (빅 데이터 환경에서 계층적 문서 유형 분류를 위한 클러스터링 기반 다중 SVM 모델)

  • Kim, Young Soo;Lee, Byoung Yup
    • The Journal of the Korea Contents Association
    • /
    • v.17 no.11
    • /
    • pp.600-608
    • /
    • 2017
  • Recently data growth rates are growing exponentially according to the rapid expansion of internet. Since users need some of all the information, they carry a heavy workload for examination and discovery of the necessary contents. Therefore information retrieval must provide hierarchical class information and the priority of examination through the evaluation of similarity on query and documents. In this paper we propose an Multi-class support vector machines model based clustering for hierarchical document categorization that make semantic search possible considering the word co-occurrence measures. A combination of hierarchical document categorization and SVM classifier gives high performance for analytical classification of web documents that increase exponentially according to extension of document hierarchy. More information retrieval systems are expected to use our proposed model in their developments and can perform a accurate and rapid information retrieval service.

A Topic Analysis of SW Education Textdata Using R (R을 활용한 SW교육 텍스트데이터 토픽분석)

  • Park, Sunju
    • Journal of The Korean Association of Information Education
    • /
    • v.19 no.4
    • /
    • pp.517-524
    • /
    • 2015
  • In this paper, to find out the direction of interest related to the SW education, SW education news data were gathered and its contents were analyzed. The topic analysis of SW education news was performed by collecting the data of July 23, 2013 to October 19, 2015. By analyzing the relationship among the most mentioned top 20 words with the web crawling using R, the result indicated that the 20 words are the closely relevant data as the thickness of the node size of the 20 words was balancing each other in the co-occurrence matrix graph focusing on the 'SW education' word. Moreover, our analysis revealed that the data were mainly composed of the topics about SW talent, SW support Program, SW educational mandate, SW camp, SW industry and the job creation. This could be used for big data analysis to find out the thoughts and interests of such people in the SW education.

Analysis on Topics of Digital Preservation Researches and Courses (디지털 보존 관련 학술연구 및 교과 주제분석)

  • Jeong, Uiyeon;Choi, Sanghee
    • Journal of the Korean Society for Library and Information Science
    • /
    • v.53 no.3
    • /
    • pp.25-43
    • /
    • 2019
  • Recently there has been a growing interest in digital preservation and digital curation with rapid increase of digital resource. This study aims to investigate the research topics and the course topics related digital preservation and digital curation. The course information is collected from the curricular of library and information science departments and archival science departments in leading countries such as US, England, Ireland, Canada and New Zealand. Title keyword profiling and network analysis were adapted to discover core research and education areas. The key topics in the abstracts of research papers and the contents of the course were also illustrated by these methods. In the research analysis, archival system is the biggest area of researches related digital preservation and digital curation. Courser analysis shows digital curation education and process is the important area of education. As a result of content analysis, plan and strategy is a notable topic of research and record management process is a major topic of courses for digital preservation and digital curation. In addition, format of digital resource is an important topic for research and courses.

Identification of Strategic Fields for Developing Smart City in Busan Using Text Mining (텍스트 마이닝을 이용한 스마트 도시계획 수립을 위한 전략분야 도출연구: 부산 사례를 바탕으로)

  • Chae, Yoonsik;Lee, Sanghoon
    • Journal of Digital Convergence
    • /
    • v.16 no.11
    • /
    • pp.1-15
    • /
    • 2018
  • The purpose of this study is to analyze bibliographic information of Busan and other cities' reports for urban development initiative and identify the strategic fields for future smart city plan. Text mining method is used in this study to extract keywords and identify the characteristics and patterns of information in urban development reports. As a result, in earlier stage, Busan city focused on service creation for industrial development but there are lack of discussions on the linkage of information systems with ICT technology. However, recent urban planning in Busan contained various contents related to integrated connections of infrastructure, ICT system, and operation management of city in the specific fields of traffic, tourism, welfare, port/logistics, culture/MICE. This results of study is expected to provide policy implications for planning the future urban initiatives of smart city development.

Network Analysis of the Intellectual Structure of Addiction Research in Social Sciences: Based on the KCI Articles Published in 2019 (사회과학 중독연구 분야의 지적구조에 관한 네트워크 분석 : 2019년도 KCI 등재 논문을 기반으로)

  • Lee, Serim;Chun, JongSerl
    • The Journal of the Korea Contents Association
    • /
    • v.21 no.10
    • /
    • pp.21-37
    • /
    • 2021
  • This study investigated the intellectual structure of the latest trends in Korean addiction research in the social sciences. A network analysis of keywords with co-word occurrence was performed on 172 papers from the KCI database based on the data from the year of 2019, and a total of 432 keywords were extracted. The network analysis was performed using several programs: Bibexcel, COOC, WNET, and NodeXL. As a result of the study, keywords related to addiction type, study subjects, research methods, and research variables were found, and a total of 20 clusters were identified. Furthermore, to identify and measure weighted networks, the relationships between each keyword were explored and discussed in detail through a network analysis of global centralities, local centralities, and betweenness centralities. The study indicated that the latest issues were focused on smartphone addiction and provided implications for the future research and practice that fields and topics of relationship addiction, food addiction, and work addiction should be more considered. Further, the study discussed the relationship between drug addiction-crime, alcohol addiction-family, and gambling addiction-motivation and the necessity of qualitative study.

The Effects of Perceived Justice on Store Loyalty in the Department Stores Service Recovery (백화점 서비스 회복과정의 지각된 공정성에 점포 애호도에 미치는 영향)

  • Kim, Yong-Han;Bae, Mu-Eun
    • Journal of Distribution Research
    • /
    • v.10 no.3
    • /
    • pp.59-86
    • /
    • 2005
  • This study examined whether the efforts of department store for recovering services may be perceived fairly from the standpoint of customers upon any occurrence of service failure, whether such perceived justice contributes to higher customer satisfaction and trust, and whether such customer satisfaction and trust have positive influence on store loyalty of customers, respectively. For this sake, this study investigated relevant literatures, set up some hypotheses to solve main questionable considerations and made a corresponding empirical analysis. For empirical analysis, a questionnaire survey was applied to total 204 customers who experienced in service recovery around domestic major department stores in the last one(1) year. With regard to empirical analysis to verify some hypotheses hereof, this study verified the reliability and validity of each questionnaire item by means of statistical programs like SPSS 10,0 and AMOS 4.0, followed by verifying hypotheses through SEM(structural equation modeling) analysis. These results indicate that positive efforts of department stores for service recovery upon their service failure may bring customer satisfaction and trust in department store, which result in boosting up store loyalty of customers. In this empirical study, it was notably demonstrated that despite any occurrence of service failure, the corresponding efforts of department stores to improve the level of justice perceived by customers in course of service recovery, had positive effects on better customer satisfaction, trust and store loyalty, and these effects induced customer's continuous repurchase, positive WOM(word of mouth) and recommendation about the use of appropriate department store, which may contribute to better its competitive edge.

  • PDF

Comparative Analysis of Low Fertility Policy and the Public Perceptions using Text-Mining Methodology (텍스트 마이닝을 활용한 저출산 정책과 대중인식 비교)

  • Bae, Giryeon;Moon, HyunJeong;Lee, Jaeil;Park, Mina;Park, Arum
    • Journal of Digital Convergence
    • /
    • v.19 no.12
    • /
    • pp.29-42
    • /
    • 2021
  • As the low fertility intensifies in Korea, this study investigated fundamental differences between the government's low fertility policy and public perception of it. To this end, we selected four times 'Aging Society and Population Policy' documents and news comments for two weeks immediately after announcement of the third and fourth Policy as analysis targets. Then we conducted word frequency analysis, co-occurrence analysis and CONCOR analysis. As a result of analyses, first, direct childcare support during the first and second periods, and a social structural approach during third and fourth periods were noticeable. Second, it was revealed that both policies and comments aim for the work-family compatibility in 'parenting'. Lastly it was showed public interest in environment of raising children and the critical mind to effectiveness of the policy. This study is meaningful in that it confirmed the public perception using big data analysis, and it will help improve the direction for the future low fertility policy.