• Title/Summary/Keyword: term co-occurrence

Search Result 53, Processing Time 0.029 seconds

Measurement of Document Similarity using Word and Word-Pair Frequencies (단어 및 단어쌍 별 빈도수를 이용한 문서간 유사도 측정)

  • 김혜숙;박상철;김수형
    • Proceedings of the IEEK Conference
    • /
    • 2003.07d
    • /
    • pp.1311-1314
    • /
    • 2003
  • In this paper, we propose a method to measure document similarity. First, we have exploited single-term method that extracts nouns by using a lexical analyzer as a preprocessing step to match one index to one noun. In spite of irrelevance between documents, possibility of increasing document similarity is high with this method. For this reason, a term-phrase method has been reported. This method constructs co-occurrence between two words as an index to measure document similarity. In this paper, we tried another method that combine these two methods to compensate the problems in these two methods. Six types of features are extracted from two input documents, and they are fed into a neural network to calculate the final value of document similarity. Reliability of our method has been proved by an experiment of document retrieval.

  • PDF

Query Term Expansion and Reweighting using Term Co-Occurrence Similarity and Fuzzy Inference (용어 발생 유사도와 퍼지 추론을 이용한 질의 용어 확장 및 가중치 재산정)

  • Kim, Ju-Yeon;Kim, Byeong-Man
    • Journal of KIISE:Software and Applications
    • /
    • v.27 no.9
    • /
    • pp.961-972
    • /
    • 2000
  • 본 논문에서는 사용자의 적합 피드백을 기반으로 적합 문서들에서 발생하는 용어들과 초기 질의어간의 발생 빈도 유사도 및 퍼지 추론을 이용하여 용어의 가중치를 산정하는 방법에 대하여 제안한다. 피드백 문서들에서 발생하는 용어들 중에서 불용어를 제외한 모든 용어들을 질의어로 확장될 수 있는 후보 용어들로 선택하고, 발생 빈도 유사성을 이용한 초기 질의어-후보 용어의 관련 정도, 용어의 IDF, DF 정보를 퍼지 추론에 적용하여 후보 용어의 초기 질의어에 대한 최종적인 관련 정도를 산정 하였으며, 피드백 문서들에서의 가중치와 관련 정도를 결합하여 후보 용어들의 가중치를 산정 하였다. 본 논문에서는 성능을 평가하기 위하여 KT-set 1.0과 KT-set 2.0을 사용하였으며, 성능의 상대적인 평가를 위하여 Dec-Hi 방법, 용어 분포 유사도를 이용한 방법, 퍼지 추론을 이용한 방법들을 정확률-재현률을 사용하여 평가하였다.

  • PDF

Pancreatic Fistula after D1+/D2 Radical Gastrectomy according to the Updated International Study Group of Pancreatic Surgery Criteria: Risk Factors and Clinical Consequences. Experience of Surgeons with High Caseloads in a Single Surgical Center in Eastern Europe

  • Martiniuc, Alexandru;Dumitrascu, Traian;Ionescu, Mihnea;Tudor, Stefan;Lacatus, Monica;Herlea, Vlad;Vasilescu, Catalin
    • Journal of Gastric Cancer
    • /
    • v.21 no.1
    • /
    • pp.16-29
    • /
    • 2021
  • Purpose: Incidence, risk factors, and clinical consequences of pancreatic fistula (POPF) after D1+/D2 radical gastrectomy have not been well investigated in Western patients, particularly those from Eastern Europe. Materials and Methods: A total of 358 D1+/D2 radical gastrectomies were performed by surgeons with high caseloads in a single surgical center from 2002 to 2017. A retrospective analysis of data that were prospectively gathered in an electronic database was performed. POPF was defined and graded according to the International Study Group for Pancreatic Surgery (ISGPS) criteria. Uni- and multivariate analyses were performed to identify potential predictors of POPF. Additionally, the impact of POPF on early complications and long-term outcomes were investigated. Results: POPF was observed in 20 patients (5.6%), according to the updated ISGPS grading system. Cardiovascular comorbidities emerged as the single independent predictor of POPF formation (risk ratio, 3.051; 95% confidence interval, 1.161-8.019; P=0.024). POPF occurrence was associated with statistically significant increased rates of postoperative hemorrhage requiring re-laparotomy (P=0.029), anastomotic leak (P=0.002), 90-day mortality (P=0.036), and prolonged hospital stay (P<0.001). The long-term survival of patients with gastric adenocarcinoma was not affected by POPF (P=0.661). Conclusions: In this large series of Eastern European patients, the clinically relevant rate of POPF after D1+/D2 radical gastrectomy was low. The presence of co-existing cardiovascular disease favored the occurrence of POPF and was associated with an increased risk of postoperative bleeding, anastomotic leak, 90-day mortality, and prolonged hospital stay. POPF was not found to affect the long-term survival of patients with gastric adenocarcinoma.

A Semantic Representation Based-on Term Co-occurrence Network and Graph Kernel

  • Noh, Tae-Gil;Park, Seong-Bae;Lee, Sang-Jo
    • International Journal of Fuzzy Logic and Intelligent Systems
    • /
    • v.11 no.4
    • /
    • pp.238-246
    • /
    • 2011
  • This paper proposes a new semantic representation and its associated similarity measure. The representation expresses textual context observed in a context of a certain term as a network where nodes are terms and edges are the number of cooccurrences between connected terms. To compare terms represented in networks, a graph kernel is adopted as a similarity measure. The proposed representation has two notable merits compared with previous semantic representations. First, it can process polysemous words in a better way than a vector representation. A network of a polysemous term is regarded as a combination of sub-networks that represent senses and the appropriate sub-network is identified by context before compared by the kernel. Second, the representation permits not only words but also senses or contexts to be represented directly from corresponding set of terms. The validity of the representation and its similarity measure is evaluated with two tasks: synonym test and unsupervised word sense disambiguation. The method performed well and could compete with the state-of-the-art unsupervised methods.

Towards Next Generation Multimedia Information Retrieval by Analyzing User-centered Image Access and Use (이용자 중심의 이미지 접근과 이용 분석을 통한 차세대 멀티미디어 검색 패러다임 요소에 관한 연구)

  • Chung, EunKyung
    • Journal of the Korean Society for Library and Information Science
    • /
    • v.51 no.4
    • /
    • pp.121-138
    • /
    • 2017
  • As information users seek multimedia with a wide variety of information needs, information environments for multimedia have been developed drastically. More specifically, as seeking multimedia with emotional access points has been popular, the needs for indexing in terms of abstract concepts including emotions have grown. This study aims to analyze the index terms extracted from Getty Image Bank. Five basic emotion terms, which are sadness, love, horror, happiness, anger, were used when collected the indexing terms. A total 22,675 index terms were used for this study. The data are three sets; entire emotion, positive emotion, and negative emotion. For these three data sets, co-word occurrence matrices were created and visualized in weighted network with PNNC clusters. The entire emotion network demonstrates three clusters and 20 sub-clusters. On the other hand, positive emotion network and negative emotion network show 10 clusters, respectively. The results point out three elements for next generation of multimedia retrieval: (1) the analysis on index terms for emotions shown in people on image, (2) the relationship between connotative term and denotative term and possibility for inferring connotative terms from denotative terms using the relationship, and (3) the significance of thesaurus on connotative term in order to expand related terms or synonyms for better access points.

Design of WWW IR System Based on Keyword Clustering Architecture (색인어 말뭉치 처리를 기반으로 한 웹 정보검색 시스템의 설계)

  • 송점동;이정현;최준혁
    • The Journal of Information Technology
    • /
    • v.1 no.1
    • /
    • pp.13-26
    • /
    • 1998
  • In general Information retrieval systems, improper keywords are often extracted and different search results are offered comparing to user's aim bacause the systems use only term frequency informations for selecting keywords and don't consider their meanings. It represents that improving precision is limited without considering semantics of keywords because recall ratio and precision have inverse proportion relation. In this paper, a system which is able to improve precision without decreasing recall ratio is designed and implemented, as client user module is introduced which can send feedbacks to server with user's intention. For this purpose, keywords are selected using relative term frequency and inverse document frequency and co-occurrence words are extracted from original documents. Then, the keywords are clustered by their semantics using calculated mutual informations. In this paper, the system can reject inappropriate documents using segmented semantic informations according to feedbacks from client user module. Consequently precision of the system is improved without decreasing recall ratio.

  • PDF

Document Summarization Based on Sentence Clustering Using Graph Division (그래프 분할을 이용한 문장 클러스터링 기반 문서요약)

  • Lee Il-Joo;Kim Min-Koo
    • The KIPS Transactions:PartB
    • /
    • v.13B no.2 s.105
    • /
    • pp.149-154
    • /
    • 2006
  • The main purpose of document summarization is to reduce the complexity of documents that are consisted of sub-themes. Also it is to create summarization which includes the sub-themes. This paper proposes a summarization system which could extract any salient sentences in accordance with sub-themes by using graph division. A document can be represented in graphs by using chosen representative terms through term relativity analysis based on co-occurrence information. This graph, then, is subdivided to represent sub-themes through connected information. The divided graphs are types of sentence clustering which shows a close relationship. When salient sentences are extracted from the divided graphs, summarization consisted of core elements of sentences from the sub-themes can be produced. As a result, the summarization quality will be improved.

Topic Modeling Analysis of Social Media Marketing using BERTopic and LDA

  • YANG, Woo-Ryeong;YANG, Hoe-Chang
    • The Journal of Industrial Distribution & Business
    • /
    • v.13 no.9
    • /
    • pp.37-50
    • /
    • 2022
  • Purpose: The purpose of this study is to explore and compare research trends in Korea and overseas academic papers on social media marketing, and to present new academic perspectives for the future direction in Korea. Research design, data and methodology: We used English abstract of research paper (Korea's: 1,349, overseas': 5,036) for word frequency analysis, topic modeling, and trend analysis for each topic. Results: The results of word frequency and co-occurrence frequency analysis showed that Korea researches focused on the experiential values of users, and overseas researches focused on platforms and content. Next, 13 topics and 12 topics for Korea and overseas researches were derived from topic modeling. And, trend analysis showed that Korean studies were different from overseas in applying marketing methods to specific industries and they were interested in the short-term performance of social media marketing. Conclusions: We found that the long-term strategies of social media marketing and academic interest in the overall industry will necessary in the future researches. Also, data mining techniques will necessary to generate more general results by quantifying various phenomena in reality. Finally, we expected that continuous and various academic approaches for volatile social media is effective to derive practical implications.

A Technical Approach for Suggesting Research Directions in Telecommunications Policy

  • Oh, Junseok;Lee, Bong Gyou
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • v.8 no.12
    • /
    • pp.4467-4488
    • /
    • 2014
  • The bibliometric analysis is widely used for understanding research domains, trends, and knowledge structures in a particular field. The analysis has majorly been used in the field of information science, and it is currently applied to other academic fields. This paper describes the analysis of academic literatures for classifying research domains and for suggesting empty research areas in the telecommunications policy. The application software is developed for retrieving Thomson Reuters' Web of Knowledge (WoK) data via web services. It also used for conducting text mining analysis from contents and citations of publications. We used three text mining techniques: the Keyword Extraction Algorithm (KEA) analysis, the co-occurrence analysis, and the citation analysis. Also, R software is used for visualizing the term frequencies and the co-occurrence network among publications. We found that policies related to social communication services, the distribution of telecommunications infrastructures, and more practical and data-driven analysis researches are conducted in a recent decade. The citation analysis results presented that the publications are generally received citations, but most of them did not receive high citations in the telecommunications policy. However, although recent publications did not receive high citations, the productivity of papers in terms of citations was increased in recent ten years compared to the researches before 2004. Also, the distribution methods of infrastructures, and the inequity and gap appeared as topics in important references. We proposed the necessity of new research domains since the analysis results implies that the decrease of political approaches for technical problems is an issue in past researches. Also, insufficient researches on policies for new technologies exist in the field of telecommunications. This research is significant in regard to the first bibliometric analysis with abstracts and citation data in telecommunications as well as the development of software which has functions of web services and text mining techniques. Further research will be conducted with Big Data techniques and more text mining techniques.

Effects of Postharvest Treatment of Plastic Film, Ethylene Scrubber, and Prolong on the Market Quality in 'Niitaka' Pears during Storage and Simulated Marketing (동양(東洋) 배 '신고(新高)'의 저장전(貯藏前) Plastic Film, Ethylene 제거제(除去劑) 및 Prolong처리(處理)가 저장(貯藏)과 유통조건(流通條件)에서 상품성(商品性)에 미치는 영향(影響))

  • Lee, Jae Chang;Hwang, Yong Soo
    • Korean Journal of Agricultural Science
    • /
    • v.19 no.2
    • /
    • pp.145-152
    • /
    • 1992
  • This experiment was planned to find a proper postharvest handling technique of 'Niitaka' pears for storage. The effect of polyethylene film wrapping, ethylene scrubber, and Prolong application on maintanence of freshness were compared in prestored fruit.(60 days at $0^{\circ}C$). Weight loss was confirmed to be a major factor responsible for freshness loss during storage. Polyethylene film wrapping greatly reduced weight loss during storage but increased light skin browning. Also, in long-term storage, polyethylene film wrapping appeared not to be appropriate due to the severe occurrence of tissue senescence and/or senescence breakdown. Prolong application was found not to be effective on reducing weight loss as well as keeping freshness. Ethylene scrubber in polyethylene film wrapping effectively reduced the occurrence of light skin browning.

  • PDF