Search | Korea Science

BERT Sparse: Keyword-based Document Retrieval using BERT in Real time (BERT Sparse: BERT를 활용한 키워드 기반 실시간 문서 검색)

Kim, Youngmin;Lim, Seungyoung;Yu, Inguk;Park, Soyoon
- Annual Conference on Human and Language Technology
- /
- 2020.10a
- /
- pp.3-8
- /
- 2020
문서 검색은 오래 연구되어 온 자연어 처리의 중요한 분야 중 하나이다. 기존의 키워드 기반 검색 알고리즘 중 하나인 BM25는 성능에 명확한 한계가 있고, 딥러닝을 활용한 의미 기반 검색 알고리즘의 경우 문서가 압축되어 벡터로 변환되는 과정에서 정보의 손실이 생기는 문제가 있다. 이에 우리는 BERT Sparse라는 새로운 문서 검색 모델을 제안한다. BERT Sparse는 쿼리에 포함된 키워드를 활용하여 문서를 매칭하지만, 문서를 인코딩할 때는 BERT를 활용하여 쿼리의 문맥과 의미까지 반영할 수 있도록 고안하여, 기존 키워드 기반 검색 알고리즘의 한계를 극복하고자 하였다. BERT Sparse의 검색 속도는 BM25와 같은 키워드 기반 모델과 유사하여 실시간 서비스가 가능한 수준이며, 성능은 Recall@5 기준 93.87%로, BM25 알고리즘 검색 성능 대비 19% 뛰어나다. 최종적으로 BERT Sparse를 MRC 모델과 결합하여 open domain QA환경에서도 F1 score 81.87%를 얻었다.
PDF

A study on Similarity analysis of National R&D Programs using R&D Project's technical classification (R&D과제의 기술분류를 이용한 사업간 유사도 분석 기법에 관한 연구)

Kim, Ju-Ho;Kim, Young-Ja;Kim, Jong-Bae
- Journal of Digital Contents Society
- /
- v.13 no.3
- /
- pp.317-324
- /
- 2012
Recently, coordination task of similarity between national R&D programs is emphasized on view from the R&D investment efficiency. But the previous similarity search method like text-based similarity search which using keyword of R&D projects has reached the limit due to deviation of document's quality. For the solve the limitations of text-based similarity search using the keyword extraction, in this study, utilization of R&D project's technical classification will be discussed as a new similarity search method when analyzed of similarity between national R&D programs. To this end, extracts the Science and Technology Standard Classification of R & D projects which are collected when national R&D Survey & analysis, and creates peculiar vector model of each R&D programs. Verify a reliability of this study by calculate the cosine-based and Euclidean distance-based similarity and compare with calculated the text-based similarity.
https://doi.org/10.9728/dcs.2012.13.3.317 인용 PDF KSCI

Automatic Construction of Reduced Dimensional Cluster-based Keyword Association Networks using LSI (LSI를 이용한 차원 축소 클러스터 기반 키워드 연관망 자동 구축 기법)

Yoo, Han-mook;Kim, Han-joon;Chang, Jae-young
- Journal of KIISE
- /
- v.44 no.11
- /
- pp.1236-1243
- /
- 2017
In this paper, we propose a novel way of producing keyword networks, named LSI-based ClusterTextRank, which extracts significant key words from a set of clusters with a mutual information metric, and constructs an association network using latent semantic indexing (LSI). The proposed method reduces the dimension of documents through LSI, decomposes documents into multiple clusters through k-means clustering, and expresses the words within each cluster as a maximal spanning tree graph. The significant key words are identified by evaluating their mutual information within clusters. Then, the method calculates the similarities between the extracted key words using the term-concept matrix, and the results are represented as a keyword association network. To evaluate the performance of the proposed method, we used travel-related blog data and showed that the proposed method outperforms the existing TextRank algorithm by about 14% in terms of accuracy.
https://doi.org/10.5626/JOK.2017.44.11.1236 인용 KSCI

An Improvement study in Keyword-centralized academic information service - Based on Recommendation and Classification in NDSL - (키워드 중심 학술정보서비스 개선 연구 - NDSL 추천 및 분류를 중심으로 -)

Kim, Sun-Kyum;Kim, Wan-Jong;Lee, Tae-Seok;Bae, Su-Yeong
- Journal of Korean Library and Information Science Society
- /
- v.49 no.4
- /
- pp.265-294
- /
- 2018
Nowadays, due to an explosive increase in information, information filtering is very important to provide proper information for users. Users hardly obtain scholarly information from a huge amount of information in NDSL of KISTI, except for simple search. In this paper, we propose the service, PIN to solve this problem. Pin provides the word cloud including analyzed users' and others' interesting, co-occurence, and searched keywords, rather than the existing word cloud simply consisting of all keywords and so offers user-customized papers, reports, patents, and trends. In addition, PIN gives the paper classification in NDSL according to keyword matching based classification with the overlapping classification enabled-academic classification system for better search and access to solve this problem. In this paper, Keywords are extracted according to the classification from papers published in Korean journals in 2016 to design classification model and we verify this model.
https://doi.org/10.16981/kliss.49.201812.265 인용 PDF KSCI

Tag Search System Using the Keyword Extraction and Similarity Evaluation (키워드 추출 및 유사도 평가를 통한 태그 검색 시스템)

Jung, Jaein;Yoo, Myungsik
- The Journal of Korean Institute of Communications and Information Sciences
- /
- v.40 no.12
- /
- pp.2485-2487
- /
- 2015
Recently, Hashtag is widely used in SNS like Facebook, Twitter and personal blogs. However, the efficiency of tag search system is poor due to the indiscriminate use of hashtags. To enhance the accuracy of tag search system, we proposed a tag search system using the keyword extraction and similarity evaluation. The experimental results show that the proposed system provides the higher accuracy on tag search results.
https://doi.org/10.7840/kics.2015.40.12.2485 인용 PDF KSCI

Design of Document Suggestion System based on TF-IDF Algorithm for Efficient Organization of Documentation (효율적인 문서 구성을 위한 TF-IDF 알고리즘 기반 문서 제안 시스템의 설계)

Kim, Young-Hoon;Park, Seung-Min;Cho, Dae-Soo
- Proceedings of the Korean Society of Computer Information Conference
- /
- 2022.07a
- /
- pp.527-528
- /
- 2022
빠르게 변하는 환경에 맞춰 평생 교육이 일반화되고 개인에게 요구되는 학습량은 많아지고 있으며 높아진 학습량에 맞게 학습 시간 단축과 효율적인 학습을 위한 학습 방법을 선택하는 것이 중요해지고 있다. 본 논문에서는 학습 정리를 위해 작성한 문서를 분석하여 해당 문서와 관련된 문서를 제안하고 본 문서와 엮어 학습을 위한 문서 묶음을 만들 수 있는 시스템을 제안한다. 문서의 유사도, 중요도를 구할 수 있는 TF-IDF를 이용하여 문서를 분석해 키워드를 추출한 다음 그와 관련된 문서를 제안하고 문서 묶음을 만들어 조회할 수 있도록 한다. 이 시스템은 학습 정리 시 관련 문서를 함께 볼 수 있도록 하고, 필요하다면 묶음으로 만들어 효과적인 학습을 위한 도구로 이용할 수 있다.
PDF

A System for Supporting Lyrics Writing Using Lyrics Data (가사 데이터 기반의 작사 지원 시스템 연구)

Young-Jae Park;Heeryon Cho
- Proceedings of the Korea Information Processing Society Conference
- /
- 2023.05a
- /
- pp.351-352
- /
- 2023
본 논문은 과거 한국 가요(K 팝)의 가사를 수집하여 (1) 특정 키워드와 관련된 기존 가사를 검색하거나, (2) 작사가가 작성한 새로운 가사와 유사한 기존 가사를 검색하거나, (3) 특정 키워드와 관련된 가사 속 어휘를 제안하는 작사 지원 시스템을 제안한다. 지금까지의 음악 관련 시스템은 음악을 소비하는 사람들을 위한 음악 추천 시스템에 집중해 왔으나, 이 연구에서는 음악을 생산하는 작사가에게 초점을 맞춰 이들을 돕는 작사 지원 시스템을 제안하고자 한다. 제안 시스템은 TF-IDF 와 word2vec 을 활용하여 가사와 단어 벡터 공간에 가사와 어휘를 배치하고 코사인 유사도를 계산한다.
https://doi.org/10.3745/PKIPS.y2023m05a.351 인용 PDF

Enhancing Classification Performance of Temporal Keyword Data by Using Moving Average-based Dynamic Time Warping Method (이동 평균 기반 동적 시간 와핑 기법을 이용한 시계열 키워드 데이터의 분류 성능 개선 방안)

Jeong, Do-Heon
- Journal of the Korean Society for information Management
- /
- v.36 no.4
- /
- pp.83-105
- /
- 2019
This study aims to suggest an effective method for the automatic classification of keywords with similar patterns by calculating pattern similarity of temporal data. For this, large scale news on the Web were collected and time series data composed of 120 time segments were built. To make training data set for the performance test of the proposed model, 440 representative keywords were manually classified according to 8 types of trend. This study introduces a Dynamic Time Warping(DTW) method which have been commonly used in the field of time series analytics, and proposes an application model, MA-DTW based on a Moving Average(MA) method which gives a good explanation on a tendency of trend curve. As a result of the automatic classification by a k-Nearest Neighbor(kNN) algorithm, Euclidean Distance(ED) and DTW showed 48.2% and 66.6% of maximum micro-averaged F1 score respectively, whereas the proposed model represented 74.3% of the best micro-averaged F1 score. In all respect of the comprehensive experiments, the suggested model outperformed the methods of ED and DTW.
https://doi.org/10.3743/KOSIM.2019.36.4.083 인용 PDF KSCI

Improving Diversity of Keyword Search on Graph-structured Data by Controlling Similarity of Content Nodes (콘텐트 노드의 유사성 제어를 통한 그래프 구조 데이터 검색의 다양성 향상)

Park, Chang-Sup
- The Journal of the Korea Contents Association
- /
- v.20 no.3
- /
- pp.18-30
- /
- 2020
Recently, as graph-structured data is widely used in various fields such as social networks and semantic Webs, needs for an effective and efficient search on a large amount of graph data have been increasing. Previous keyword-based search methods often find results by considering only the relevance to a given query. However, they are likely to produce semantically similar results by selecting answers which have high query relevance but share the same content nodes. To improve the diversity of search results, we propose a top-k search method that finds a set of subtrees which are not only relevant but also diverse in terms of the content nodes by controlling their similarity. We define a criterion for a set of diverse answer trees and design two kinds of diversified top-k search algorithms which are based on incremental enumeration and A^⁎ heuristic search, respectively. We also suggest an improvement on the A^⁎ search algorithm to enhance its performance. We show by experiments using real data sets that the proposed heuristic search method can find relevant answers with diverse content nodes efficiently.
https://doi.org/10.5392/JKCA.2020.20.03.018 인용 PDF KSCI HTML

Folder Recommendation Based on User Knowledge (사용자 지식을 반영한 메일 폴더 추천 방법론)

You Mee;Park Joo Seok;Kim Jae Kyeong
- Journal of Intelligence and Information Systems
- /
- v.10 no.3
- /
- pp.133-146
- /
- 2004
By the development of the network technology, the types and amount of information that users keep in contact with have been dramatically increased. As a result, users are consuming a lot of time and energy to find needed information. On this, this article presents a new methodology that can efficiently manage their information within small cost by using content-based recommendation method and keyword affinity method. By using keyword affinity method, this methodology solves the content-based recommendation method's weak point that the performance is not good within the environment that the preferences of users are rapidly changing and new contents are created continuously and the accuracy level is low until the information of preferences are sufficiently gathered. This article carried out research on the personal e-mail environment where new information is frequently created and disappeared. Also this article assists folder recommendation for the efficient management of e-mail and verified the methodology mentioned above by an experiment to compare the performance of existing folder recommendation methods with the performance of this new method.
PDF

Search Result 311, Processing Time 0.023 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)