• 제목/요약/키워드: term co-occurrence

검색결과 53건 처리시간 0.022초

단어의 공기정보를 이용한 클러스터 기반 다중문서 요약 (Multi-document Summarization Based on Cluster using Term Co-occurrence)

  • 이일주;김민구
    • 한국정보과학회논문지:소프트웨어및응용
    • /
    • 제33권2호
    • /
    • pp.243-251
    • /
    • 2006
  • 대표문장 추출에 의한 다중문서 요약에서는 비슷한 정보가 여러 문서에서 반복적으로 나타나는 정보의 중복문제에 대해 문장의 유사성과 차이점을 고려하여 이를 해결할 수 있는 효율적인 방법이 필요하다. 본 논문에서는 단어의 공기정보에 의한 관련단어 클러스터링 기법을 이용하여 문장의 중복성을 제거하고 중요문장을 추출하는 다중문서 요약을 제안한다. 관련단어 클러스터링 기법에서는 각 단어들은 서로 독립적으로 존재하는 것이 아니라 서로 간에 의미적으로 연관되어 있다고 보며 주제별 문장클러스터단위의 단어 연관성(cohesion)을 이용한다. 평가용 실험문서인 DUC(Document Understanding Conferences) 데이타를 이용하여 실험한 결과 본 논문에서 제안한 문장클러스터단위의 단어 공기정보를 이용한 방법이 단순 통계정보와 문서단위 단어 공기정보, 문장단위 단어 공기정보에 의한 다중문서 요약에 비해 좋은 결과를 보였다.

Text Mining of Wood Science Research Published in Korean and Japanese Journals

  • Eun-Suk JANG
    • Journal of the Korean Wood Science and Technology
    • /
    • 제51권6호
    • /
    • pp.458-469
    • /
    • 2023
  • Text mining techniques provide valuable insights into research information across various fields. In this study, text mining was used to identify research trends in wood science from 2012 to 2022, with a focus on representative journals published in Korea and Japan. Abstracts from Journal of the Korean Wood Science and Technology (JKWST, 785 articles) and Journal of Wood Science (JWS, 812 articles) obtained from the SCOPUS database were analyzed in terms of the word frequency (specifically, term frequency-inverse document frequency) and co-occurrence network analysis. Both journals showed a significant occurrence of words related to the physical and mechanical properties of wood. Furthermore, words related to wood species native to each country and their respective timber industries frequently appeared in both journals. CLT was a common keyword in engineering wood materials in Korea and Japan. In addition, the keywords "MDF," "MUF," and "GFRP" were ranked in the top 50 in Korea. Research on wood anatomy was inferred to be more active in Japan than in Korea. Co-occurrence network analysis showed that words related to the physical and structural characteristics of wood were organically related to wood materials.

한국어 정보 검색에서 의미적 용어 불일치 완화 방안 (Alleviating Semantic Term Mismatches in Korean Information Retrieval)

  • 윤보현;박성진;강현규
    • 한국정보처리학회논문지
    • /
    • 제7권12호
    • /
    • pp.3874-3884
    • /
    • 2000
  • 정보검색시스템은 색인어와 질의어가 정확히 일치하지 않더라도 사용자 질의에 적합한 문서를 검색할 수 있어야 한다. 그러나, 색인어와 질의어간의 용어 불일치는 검색성능의 개선에 심각한 장애요소로 작용해 왔다. 따라서, 본 논문에서는 문서 코퍼스의 단어들간에 자동 용어 정규화를 수행하고, 용어 정규화의 산물을 한국어 정보검색 시스템에 적용하는 방안을 제시한다. 용어 불일치를 완화하기 위해 두가지 용어 정규화, 동치부류와 공기단어 클러스터를 수행한다. 첫째, 음역어, 절차오류, 그리고 동의어를 위해 문맥 유사도를 이용하여 동치부류로 구축하는 작업이다. 둘째, 상호정보와 단어 문맥의 조합을 이용하여 단어 유사도를 계산하고 문맥 기반 용어를 정규화한다. 그런 다음, K-means 알고리즘을 이용하여 자율 클러스터링을 수행하고 공기단어 클러스터를 구축한다. 본 논문에서는 이러한 용어 정규화의 산물들을 용어 불일치를 완화하기 위해 질의어 확장과정에서 사용한다. 다시 말해서 동치부류와 공기단어 클러스터는 새로운 용어로 질의를 확장하는 자원으로서 사용된다. 이러한 질의확장으로 사용자는 질의어에 음역어를 추가하여 질의어를 포괄적으로 만들거나 특정어를 추가하여 질의어를 세밀하게 만들 수 있다. 질의어 확장을 위해 두 가지 상호보완적인 방법인 용어 제시와 용어 적합성 피드백을 이용한다. 실험 결과는 제안된 시스템이 의미적 용어 불일치를 완화할 수 있고, 적절한 유사도 값을 제공할 수 있음을 보여준다. 결과적으로 제안한 시스템이 정보 검색 시스템의 검색 효율을 향상시킬 수 있음을 알 수 있다.

  • PDF

조현병 관련 주요 일간지 기사에 대한 텍스트 마이닝 분석 (Text-Mining Analyses of News Articles on Schizophrenia)

  • 남희정;류승형
    • 대한조현병학회지
    • /
    • 제23권2호
    • /
    • pp.58-64
    • /
    • 2020
  • Objectives: In this study, we conducted an exploratory analysis of the current media trends on schizophrenia using text-mining methods. Methods: First, web-crawling techniques extracted text data from 575 news articles in 10 major newspapers between 2018 and 2019, which were selected by searching "schizophrenia" in the Naver News. We had developed document-term matrix (DTM) and/or term-document matrix (TDM) through pre-processing techniques. Through the use of DTM and TDM, frequency analysis, co-occurrence network analysis, and topic model analysis were conducted. Results: Frequency analysis showed that keywords such as "police," "mental illness," "admission," "patient," "crime," "apartment," "lethal weapon," "treatment," "Jinju," and "residents" were frequently mentioned in news articles on schizophrenia. Within the article text, many of these keywords were highly correlated with the term "schizophrenia" and were also interconnected with each other in the co-occurrence network. The latent Dirichlet allocation model presented 10 topics comprising a combination of keywords: "police-Jinju," "hospital-admission," "research-finding," "care-center," "schizophrenia-symptom," "society-issue," "family-mind," "woman-school," and "disabled-facilities." Conclusion: The results of the present study highlight that in recent years, the media has been reporting violence in patients with schizophrenia, thereby raising an important issue of hospitalization and community management of patients with schizophrenia.

Influence of Microbial Activity on the Long-Term Alteration of Compacted Bentonite/Metal Chip Blocks

  • Lee, Seung Yeop;Lee, Jae-Kwang;Kwon, Jang-Soon
    • 방사성폐기물학회지
    • /
    • 제19권4호
    • /
    • pp.469-477
    • /
    • 2021
  • Safe storage of spent nuclear fuel in deep underground repositories necessitates an understanding of the long-term alteration of metal canisters and buffer materials. A small-scale laboratory alteration test was performed on metal (Cu or Fe) chips embedded in compacted bentonite blocks placed in anaerobic water for 1 year. Lactate, sulfate, and bacteria were separately added to the water to promote biochemical reactions in the system. The bentonite blocks immersed in the water were dismantled after 1 year, showing that their alteration was insignificant. However, the Cu chip exhibited some microscopic etch pits on its surface, wherein a slight sulfur component was detected. Overall, the Fe chip was more corroded than the Cu chip under the same conditions. The secondary phase of the Fe chip was locally found as carbonate materials, such as siderite (FeCO3) and calcite ((Ca, Fe)CO3). These secondary products can imply that the local carbonate occurrence on the Fe chip may be initiated and developed by an evolution (alteration) of bentonite and a diffusive provision of biogenic CO2 gas. These laboratory scale results suggest that the actual long-term alteration of metal canisters/bentonite blocks in the engineered barrier could be possible by microbial activities.

문헌정보학의 지식 구조에 관한 연구 (A Study on Intellectual Structure of Library and Information Science in Korea)

  • 유영준
    • 정보관리학회지
    • /
    • 제20권3호
    • /
    • pp.277-297
    • /
    • 2003
  • 이 연구는 색인어가 특정 주제 영역의 지식 구조를 표현할 수 있다는 것을 전제로 한다. 여기에서는 문헌정보학 관련 학술지인 정보관리학회지, 한국도서관정보학회지, 한국문헌정보학회지 등에 수록된 논문을 대상으로 국회도서관이 배정한 색인어를 클러스터링하여 문헌정보학의 지식 구조를 파악하였다. 그 과정에서, 색인어간의 연관도 및 동시 출현 빈도를 이용하여 색인어 군집을 생성하였고, 초출색인어와 시기 구분에 의한 시계열 분석을 수행함으로써 문헌정보학의 발전 과정과 그 동향을 밝혔다. 또한 색인어 군집에 의해 도출된 지식 구조와 기존의 전통적인 분류체계의 지식 구조를 비교하여 두 지식 구조간의 차이를 분석하였다.

Identification of Distinct Vaginal Microbiota Signatures Contributing Toward Preterm Birth Using an Integrative Computational Approach

  • Sudeepti Kulshreshtha;Priyanka Narad;Brojen Singh;Deepak Modi;Abhishek Sengupta
    • 한국미생물·생명공학회지
    • /
    • 제51권1호
    • /
    • pp.109-123
    • /
    • 2023
  • Preterm birth (PTB) is defined as giving birth prior to the 37th week of pregnancy and is a major cause of infant mortality. Studies have indicated that the vaginal microbiota's composition and its dysbiosis, particularly during pregnancy, may play a major role in PTB. While previous research work concentrated on well-studied microorganisms such as Lactobacillus, Prevotella, Gardnerella, various other microbes, and their significance in the vaginal microbiota's stability remain unknown. Moreover, current studies have focused primarily on the relative abundances of the microbes found, without considering their interactions with other members of the vaginal microbiota. In this work, we developed a novel computational approach and performed taxonomic classification of vaginal microbiota samples stratified longitudinally (Term/PTB) to observe compositional disparities and find underexamined microbes that may be contributing to PTB. Furthermore, we carried out a correlational analysis to build a microbial co-interaction network and investigated the functional implications of the genes present in both Term and PTB samples. The co-occurrence network revealed that Lactobacillus acts in solidarity to maintain the stability of the vaginal microbiota and did not have strong co-interactions with any of the other microbes. Similarly, microbes with strong interactions with Atopobium, a well-known marker microbe of PTB, were also observed. Additionally, several genes such as PTXA, FANCM, GPX, and DUSP were found to be playing an important role in the occurrence of PTB. This study provides a novel conceptual framework revealing distinct vaginal microbiota signatures that could be potential therapeutic targets for the prevention of PTB.

Ten Year Literature on Psychological and Behavioral Interventions Against Cancer: a Terms Analysis

  • Feng, Rui;Chai, Jing;Wang, De-Bin;Xia, Yi;Cheng, Peng-Lai;Dai, Zhao-Yang
    • Asian Pacific Journal of Cancer Prevention
    • /
    • 제13권10호
    • /
    • pp.5171-5176
    • /
    • 2012
  • We here performed a systematic review of PBIC literature using terms analysis in a hope of both identifying potential trends and patterns and exploring methods leveraging traditional literature reviews in this specific area. Articles meeting inclusion criteria were retrieved from PUBMED and translated into dichotomized article records representing presence or non-presence of MeSH terms and a metric consisting of numbers of times of co-occurrence between all pairs of terms identified using a self-designed program. The occurrence of and relations among the terms were calculated and visualized using Excel2007 and UCINET respectively. A total of 1,742 terms were identified from 997 articles retrieved. Put in a descending order, the lines representing the times of term occurrence formed a typical hyperbolic curve; when plotted along the x-axis of whole MESH terms, the lines clustered within four specific regions. Comparison of term occurrence between 2002 and 2011 revealed priority changes in population and subjects (from general groups to priority groups), intervention approaches (from medicine to exercise and psychotherapy), methodology and techniques (from cohort studies to randomized controlled trials) and outcomes (from health and mental health to quality of life, depression etc.). Networks of the terms featured a number of closely linked groups of topics including method and questionnaires, therapy and outcomes, survival management, psychological assessment and intervention, behavioral intervention (individual and community oriented). Terms analysis revealed interesting trends and patterns about PBIC publications and both the analysis methods and findings have implications for future research and literature reviews.

A Study on the General Public's Perceptions of Dental Fear Using Unstructured Big Data

  • Han-A Cho;Bo-Young Park
    • 치위생과학회지
    • /
    • 제23권4호
    • /
    • pp.255-263
    • /
    • 2023
  • Background: This study used text mining techniques to determine public perceptions of dental fear, extracted keywords related to dental fear, identified the connection between the keywords, and categorized and visualized perceptions related to dental fear. Methods: Keywords in texts posted on Internet portal sites (NAVER and Google) between 1 January, 2000, and 31 December, 2022, were collected. The four stages of analysis were used to explore the keywords: frequency analysis, term frequency-inverse document frequency (TF-IDF), centrality analysis and co-occurrence analysis, and convergent correlations. Results: In the top ten keywords based on frequency analysis, the most frequently used keyword was 'treatment,' followed by 'fear,' 'dental implant,' 'conscious sedation,' 'pain,' 'dental fear,' 'comfort,' 'taking medication,' 'experience,' and 'tooth.' In the TF-IDF analysis, the top three keywords were dental implant, conscious sedation, and dental fear. The co-occurrence analysis was used to explore keywords that appear together and showed that 'fear and treatment' and 'treatment and pain' appeared the most frequently. Conclusion: Texts collected via unstructured big data were analyzed to identify general perceptions related to dental fear, and this study is valuable as a source data for understanding public perceptions of dental fear by grouping associated keywords. The results of this study will be helpful to understand dental fear and used as factors affecting oral health in the future.

Co-occurrence Network Analysis of Keywords in Geriatric Frailty

  • Kim, Youngji;Jang, Soong-nang;Lee, Jung Lim
    • 지역사회간호학회지
    • /
    • 제29권4호
    • /
    • pp.429-439
    • /
    • 2018
  • Purpose: The aim of this study is to identify core keyword of frailty research in the past 35 years to understand the structure of knowledge of frailty. Methods: 10,367 frailty articles published between 1981 and April 2016 were retrieved from Web of Science. Keywords from these articles were extracted using Bibexcel and social network analysis was conducted with the occurrence network using NetMiner program. Results: The top five keywords with a high frequency of occurrence include 'disability', 'nursing home', 'sarcopenia', 'exercise', and 'dementia'. Keywords were classified by subheadings of MeSH and the majority of them were included under the healthcare and physical dimensions. The degree centralities of the keywords were arranged in the order of 'long term care' (0.55), 'gait' (0.42), 'physical activity' (0.42), 'quality of life' (0.42), and 'physical performance' (0.38). The betweenness centralities of the keywords were listed in the order of depression' (0.32), 'quality of life' (0.28), 'home care' (0.28), 'geriatric assessment' (0.28), and 'fall' (0.27). The cluster analysis shows that the frailty research field is divided into seven clusters: aging, sarcopenia, inflammation, mortality, frailty index, older people, and physical activity. Conclusion: After reviewing previous research in the 35 years, it has been found that only physical frailty and frailty related to medicine have been emphasized. Further research in psychological, cognitive, social, and environmental frailty is needed to understand frailty in a multifaceted and integrative manner.