Search | Korea Science

Design of Keyword Extraction System Using TFIDF (TFIDF를 이용한 키워드 추출 시스템 설계)

이말례;배환국
- Korean Journal of Cognitive Science
- /
- v.13 no.1
- /
- pp.1-11
- /
- 2002
In this paper, a test was performed to determine whether words in Anchor Text were appropriate as key words. As a result of the test. there were proper words of high weighting factor, while some others did not even appear in the text. therefore, were not appropriate as key words. In order to resolve this problem. a new method was proposed to extract key words. Using the proposed method, inappropriate key words can be removed so that new key words be set, and then, ranking becomes possible with the TFIDF value as a weighting factor of the key word. It was verified that the new method has higher accuracy compared to the previous methods.
PDF

A Study on Graph-based Topic Extraction from Microblogs (마이크로블로그를 통한 그래프 기반의 토픽 추출에 관한 연구)

Choi, Don-Jung;Lee, Sung-Woo;Kim, Jae-Kwang;Lee, Jee-Hyong
- Journal of the Korean Institute of Intelligent Systems
- /
- v.21 no.5
- /
- pp.564-568
- /
- 2011
Microblogs became popular information delivery ways due to the spread of smart phones. They have the characteristic of reflecting the interests of users more quickly than other medium. Particularly, in case of the subject which attracts many users, microblogs can supply rich information originated from various information sources. Nevertheless, it has been considered as a hard problem to obtain useful information from microblogs because too much noises are in them. So far, various methods are proposed to extract and track some subjects from particular documents, yet these methods do not work effectively in case of microblogs which consist of short phrases. In this paper, we propose a graph-based topic extraction and partitioning method to understand interests of users about a certain keyword. The proposed method contains the process of generating a keyword graph using the co-occurrences of terms in the microblogs, and the process of splitting the graph by using a network partitioning method. When we applied the proposed method on some keywords. our method shows good performance for finding a topic about the keyword and partitioning the topic into sub-topics.
https://doi.org/10.5391/JKIIS.2011.21.5.564 인용 PDF KSCI

Scientometric Analysis through Centrality Analysis of Graph for Linkage Relation of Keyword for Elder's Rehabilitation and Healthcare (노인 재활 헬스케어에 대한 키워드 연결 관계의 그래프 중심성 분석을 통한 계량 정보 분석)

Kim, Myung-Mi
- The Journal of the Korea institute of electronic communication sciences
- /
- v.14 no.2
- /
- pp.447-452
- /
- 2019
The elder problem is very serious stage in the age of present. This paper carries out scientometric analysis based on keyword that effort of global researcher for this research field as viewpoint of ICT and healthcare in order to settle physical exercise and rehabilitation of elder. First, this paper performs analysis of linkage relation of keyword. Second this paper carries out the analysis of degree distribution and centrality analysis of network based on betweenness centrality, closeness centrality and harmony centrality. Through this process, this paper reviews research trend of present and future through core keyword in the field of ICT and healthcare.
https://doi.org/10.13067/JKIECS.2019.14.2.447 인용 PDF KSCI HTML

Improving Diversity of Keyword Search on Graph-structured Data by Controlling Similarity of Content Nodes (콘텐트 노드의 유사성 제어를 통한 그래프 구조 데이터 검색의 다양성 향상)

Park, Chang-Sup
- The Journal of the Korea Contents Association
- /
- v.20 no.3
- /
- pp.18-30
- /
- 2020
Recently, as graph-structured data is widely used in various fields such as social networks and semantic Webs, needs for an effective and efficient search on a large amount of graph data have been increasing. Previous keyword-based search methods often find results by considering only the relevance to a given query. However, they are likely to produce semantically similar results by selecting answers which have high query relevance but share the same content nodes. To improve the diversity of search results, we propose a top-k search method that finds a set of subtrees which are not only relevant but also diverse in terms of the content nodes by controlling their similarity. We define a criterion for a set of diverse answer trees and design two kinds of diversified top-k search algorithms which are based on incremental enumeration and A^⁎ heuristic search, respectively. We also suggest an improvement on the A^⁎ search algorithm to enhance its performance. We show by experiments using real data sets that the proposed heuristic search method can find relevant answers with diverse content nodes efficiently.
https://doi.org/10.5392/JKCA.2020.20.03.018 인용 PDF KSCI HTML

A Implementation of Keyword Extraction Algorithm Using Anchor Text for Web's Conceptual Knowledge (웹의 개념지식을 위한 Anchor Text에서의 키워드 추출 알고리즘의 구현)

조남덕;배환국;김기태
- Proceedings of the Korean Information Science Society Conference
- /
- 2000.10b
- /
- pp.72-74
- /
- 2000
인터넷을 효과적으로 검색하기 위하여 검색엔진을 많이 이용하고 있다. 그런데 문서의 키워드를 추출할 적에 지금까지는 Anchor Text를 염두에 두지 않았었다. Anchor Text는 사람이 직접 요약한 것이고(요약성), 하이퍼링크를 포함하는 웹 문서에 반드시 존재하므로(보편성) 그 하이퍼링크가 가리키는 곳의 문서의 키워드를 추출에 적합한 용도가 될 수 있다. 웹 그래프는 이러한 Anchor Text를 이용하여 키워드를 추출함으로써 문서와 문서간, 단어와 단어간의 관계(연관성)까지도 나타내 줄 수 있게 한 검색 엔진 시스템이다. 그러나 Anchor Text 자체가 본문의 내용이 아니고, Anchor Text를 작성한 사람에 따라 다르게 작성되며, 본문의 내용과 무관한 내용도 작성할 수 있다. 따라서 Anchor Text 자체를 어떠한 여과 없이 문서의 키워드로 받아들이긴 힘들다. 본 논문에서는 TFIDF를 통해 좀 더 정확성이 있는 키워드를 추출하였다.
PDF

A Keyword Trend Analysis System Using Multiple SNS Sites (다수의 SNS를 이용한 키워드 트렌드 분석 시스템)

Lee, Myung-Chul;Han, Soo-Hyun;Lee, Jae Sung
- Proceedings of the Korea Information Processing Society Conference
- /
- 2019.10a
- /
- pp.1133-1135
- /
- 2019
기업이나 정부 등의 정책 결정에 활용하기 위해, SNS에서 사용하는 키워드를 추출하여 소비자나 유권자의 관심과 선호도를 분석하는 방법이 많이 사용되고 있다. 본 논문에서는 다수의 SNS 사이트에 올린 글과 그에 대한 공감(좋아요) 댓글, 해시태그를 분석하여 관심 키워드의 트렌드를 분석할 수 있는 시스템을 제안한다. 이 시스템에서는 각각의 SNS 글을 형태소 분석하여 키워드 빈도를 측정하고 그에 대한 공감 및 해시태그의 갯수를 계산하여 일정기간 동안의 변화를 그래프로 표시하였다. 이를 통해, 여러 사이트에서의 키워드 트렌드를 한눈에 확인할 수 있도록 했다.
https://doi.org/10.3745/PKIPS.y2019m10a.1133 인용 PDF

A Survey on Graph Mining in Social Network Service (소셜 네트워크 서비스에서의 그래프 마이닝 기법에 관한 조사)

Lee, Ji-Hyeon;Park, Young-Ho
- Proceedings of the Korea Information Processing Society Conference
- /
- 2011.11a
- /
- pp.1270-1271
- /
- 2011
소셜 네트워크 서비스는 가트너에서 2011년에 이어 2012년에도 각광받을 기술의 하나로 선정된 만큼 미래 인터넷의 핵심 키워드 중 하나로도 뽑히며, 엔터테인먼트, 검색, 방송, 커머스 등의 여러 가지 서비스와 직접 연결된다. 이러한 소셜 네트워크 서비스 가운데 하이브리드형 서비스는 사용자의 정보를 관리 및 파악하여 사용자가 원하는 제품을 예측하고 추천해주고 있으며, 이를 위해 그래프 마이닝 기술을 적용하고 있다. 하지만 그래프 마이닝 기술은 아직 복잡한 그래프 구조의 데이터에서 정보를 추출하기에 제약사항들이 발생하므로 이에 대하여 많은 연구가 활발히 이루어지고 있다. 이러한 그래프 마이닝 기술을 나아가 더 발전시켜 활용하면 기존의 하이브리드형 서비스에서 사용자의 정보를 파악하여 충성도를 높여줄 뿐 아니라 기업에서의 타켓 마케팅과 원투원 마케팅을 가능하게 해주고 기존 사용자에 대한 교차 판매와 격상판매의 전략들을 도출할 수 있을 것이다.
https://doi.org/10.3745/PKIPS.y2011m11a.1270 인용 PDF

A Study on Automatic Extraction of Core Sentences from Document using Word Cooccurrence Graph (단어의 공기 관계 그래프를 이용한 문서의 핵심 문장 추출에 관한 연구)

Ryu, Je;Han, Kwang-Rok;Sohn, Seok-Won;Rim, Kee-Wook
- The Transactions of the Korea Information Processing Society
- /
- v.7 no.11
- /
- pp.3427-3437
- /
- 2000
In this paper,we propose an method of core sciences extractionusing word cooccrrence graph in order to summarize a document. For automatic extraction of core sentenees, we construct a mean cluster from word cooccurrence graph, and find insistence which corresponds a porposed of author. And then we extract keywords by using relationship between mean cluster and isistence. Finally, core senrences are sclected based on keywords and insitances. The esults are evaluated by comparing with manual extraction, and show that the extraction performance is improved about 10%.
PDF

지능형 전자상거래를 위한 온톨로지의 효율적인 생성

Kim, Tae-Seok;Yang, Jin-Hyeok;Lee, Ji-Hong;Son, Jong-Su;Jeong, In-Jeong
- Proceedings of the Korea Inteligent Information System Society Conference
- /
- 2005.11a
- /
- pp.273-279
- /
- 2005
월드와이드웹 (WWW) 기반의 전자상거래는 주로 데이터베이스를 기반으로 서비스를 제공하고 있다. 그러나 월드와이드웹 기반의 전자상거래는 단순 키워드 검색에만 의존하고 있다. 이러한 검색은 데이터베이스 자체로는 의미적인 정보를 효과적으로 처리하기에는 많은 문제점이 있다. 1999년 말에 의미적인 정보를 효과적으로 처리하기 할 수 있는 시맨틱 웹 이 제안되었다. 시맨틱 웹은 의미적인 정보를 담고 있는 지식베이스(Knowledge Bases)인 온톨로지를 기반으로 하고 있다. 그러나 온툴로지의 생성은 많은 부분을 휴리스틱에 의존하고 있기 때문에 많은 시간과 비용이 소비된다. 따라서 우리는 이와 같은 문제를 해결하기 위하여 데이터베이스에서 온톨로지를 생성하는 방법을 제안한다. 데이터베이스는 도메인을 잘 나타내고 있는 정보의 저장소이므로 데이터베이스로부터의 온톨로지 생성은 분석, 설계 등의 사전 작업이 필요하지 않아 시간과 비용의 소비를 줄 일 수 있는 장점이 있다. 우리는 데이터베이스에서 스키마를 추출, 뼈대그래프$^{1}$ 를 생성하고 개념그래프로 확장하여 도메인을 잘 나타낼 수 있는 온톨로지를 생성하는 알고리즘을 제안하고 제안된 알고리즘을 통하여 온톨로지를 생성을 함으로서 제안된 생성 방법을 검증한다. 제안한 방법으로 생성된 온톨로지는 단순 키워드 검색에서 의미적인 검색을 할 수 있는 시맨틱 웹 서비스의 기반이 되므로 의미적 검색이 가능한 전자상거래 서비스를 구축하는데 시간과 비용의 소비를 줄임으로 차세대 전자상거래의 초석이 된다.
PDF

Cluster-based keyword Ranking Technique (클러스터 기반 키워드 랭킹 기법)

Yoo, Han-mook;Kim, Han-joon
- Proceedings of the Korea Information Processing Society Conference
- /
- 2016.10a
- /
- pp.529-532
- /
- 2016
본 논문은 기존의 TextRank 알고리즘에 상호정보량 척도를 결합하여 군집 기반에서 키워드 추출하는 ClusterTextRank 기법을 제안한다. 제안 기법은 k-means 군집화 알고리즘을 이용하여 문서들을 여러 군집으로 나누고, 각 군집에 포함된 단어들을 최소신장트리 그래프로 표현한 후 이에 근거한 군집 정보량을 고려하여 키워드를 추출한다. 제안 기법의 성능을 평가하기 위해 여행 관련 블로그 데이터를 이용하였으며, 제안 기법이 기존 TextRank 알고리즘보다 키워드 추출의 정확도가 약 13% 가량 개선됨을 보인다.
https://doi.org/10.3745/PKIPS.y2016m10a.529 인용 PDF

Search Result 51, Processing Time 0.037 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)