• Title/Summary/Keyword: 텍스트 접근법

Search Result 49, Processing Time 0.03 seconds

A WordNet-based Feature Merge Method for HyperText Classification (하이퍼텍스트 문서의 자동분류를 위한 워드넷 기반 특징 합병 기법)

  • Roh, Jun-Ho;Kim, Han-Joon;Chang, Jae-Young
    • Proceedings of the Korea Information Processing Society Conference
    • /
    • 2012.11a
    • /
    • pp.406-409
    • /
    • 2012
  • 본 논문은 하이퍼텍스트 문서의 자동분류 성능을 높이기 위한 새로운 접근법을 제시한다. 하이퍼텍스트 문서는 일반 문서와 달리 하이퍼링크로 서로 연결된 구조를 가진다. 이 하이퍼링크 정보는 대상문서와 연관도가 높은 정보를 가지고 있으며, 이러한 링크 정보로부터 특징을 보다 잘 선별하기 위해서는 보다 정밀한 접근법이 필요하다. 본 논문은 단어간 의미 유사도를 기반으로 하이퍼텍스트 링크 정보를 활용한 특징 가공기법을 제안한다. 제안 기법은 하이퍼링크 문서로부터 대상문서와 연관도가 높은 특징을 추출하기 위해 단어간 유사도 함수를 사용하며, 유사도 함수는 워드넷의 상/하위어 관계를 이용한다. 그리고 추출된 특징들 중 의미적으로 비슷한 개념의 특징들을 합병함으로써 의미적으로 보다 견고한 분류 모델을 구축한다. 제안 기법을 검증하기 위해 Web-KB 문서집합을 이용하여 실험을 수행하였고 실험 결과 기존 방법보다 우수한 성능을 보였다.

A comparative study of Entity-Grid and LSA models on Korean sentence ordering (한국어 텍스트 문장정렬을 위한 개체격자 접근법과 LSA 기반 접근법의 활용연구)

  • Kim, Youngsam;Kim, Hong-Gee;Shin, Hyopil
    • Korean Journal of Cognitive Science
    • /
    • v.24 no.4
    • /
    • pp.301-321
    • /
    • 2013
  • For the task of sentence ordering, this paper attempts to utilize the Entity-Grid model, a type of entity-based modeling approach, as well as Latent Semantic analysis, which is based on vector space modeling, The task is well known as one of the fundamental tools used to measure text coherence and to enhance text generation processes. For the implementation of the Entity-Grid model, we attempt to use the syntactic roles of the nouns in the Korean text for the ordering task, and measure its impact on the result, since its contribution has been discussed in previous research. Contrary to the case of German, it shows a positive result. In order to obtain the information on the syntactic roles, we use a strategy of using Korean case-markers for the nouns. As a result, it is revealed that the cues can be helpful to measure text coherence. In addition, we compare the results with the ones of the LSA-based model, discussing the advantages and disadvantages of the models, and options for future studies.

  • PDF

Zero-shot Text Classification based on Reinforced Learning (강화학습 기반의 제로샷 텍스트 분류)

  • Zhang Songming;Inwhee Joe
    • Proceedings of the Korea Information Processing Society Conference
    • /
    • 2023.11a
    • /
    • pp.439-441
    • /
    • 2023
  • 전통적인 텍스트 분류 방법은 상당량의 라벨링된 데이터와 미리 정의된 클래스가 필요해서 그 적용성과 확장성이 제한된다. 그래서 이런 한계를 극복하기 위해 제로샷 러닝(Zero-shot Learning)이 등장했다. 텍스트 분류 분야에서 제로샷 텍스트 분류는 모델이 대상 클래스의 샘플을 미리 접하지 않고도 인스턴스를 분류할 수 있도록 하는 중요한 주제이다. 이 문제를 해결하기 위해 정책 네트워크를 활용한 심층 강화 학습(DRL) 기반 접근법을 제안한다. 이러한 방법을 통해 모델이 새로운 의미 공간에 효과적으로 적응하면서, 다른 모델들과 비교하여 제로샷 텍스트 분류의 정확도를 향상시킬 수 있었다. XLM-R 과 비교하면 최대 15.9%의 정확도 향상이 나타났다.

Case Analysis of Bible Visualization based on Text Data Traits -Focused on Content, Structure, Quotation of Text- (텍스트 데이터의 특성에 따른 성경 시각화 사례 분석 -텍스트의 내용적, 구조적 특성 및 인용 정보를 중심으로-)

  • Kim, Hyoyoung;Park, Jin Wan
    • The Journal of the Korea Contents Association
    • /
    • v.13 no.8
    • /
    • pp.83-92
    • /
    • 2013
  • Text visualization begins with understanding text itself which is material of visual expression. To visualize any text data, sufficient understanding about characteristics of the text first and the expressive approaches can be decided depending on the derived unique characteristics of the text. In this research we aimed to establish theoretical foundation about the approaches for text visualization by diverse examples of text visualization which are derived through the various characteristics of the text. To do this, we chose the 'Bible' text which is well known globally and digital data of it can be accessed easily and thus diverse text visualization examples exist and analyzed the examples of the bible text visualization. We derived the unique characteristics of text-content, structure, quotation- as criteria for analyzing and supported validity of analysis by adopting at least 2-3 examples for each criterion. In the result, we can comprehend that the goals and expressive approaches are decided depending on the unique characteristics of the Bible text. We expect to build theoretical method for choosing the materials and approaches by analyzing more diverse examples with various point of views on the basis of this research.

Survey on Multifaceted Role of Shipping Industry and Measures to Improve Public Perception (해운산업의 다면적 역할에 대한 인식조사 및 국민인식 제고방안)

  • Lee, DongHyon
    • Journal of Korea Port Economic Association
    • /
    • v.28 no.3
    • /
    • pp.127-150
    • /
    • 2012
  • A survey showed that the public perception of the shipping industry's overall image and economic role was relatively positive. The survey revealed that public perception was mixed with respect to the multifaceted role of the shipping industry. Based on the results of this survey, this paper proposes three approaches to improving the public perception of the shipping industry. The organization contact approach includes establishing shipping institutes for city people, holding various events targeted at the public, establishing a shipping memorial hall, developing a shipping-related culture and tourism, reinventing the image of the shipping industry through a shipping-culture movement, and creating new views of the shipping industry by conducting formal education. The goods and services contact approach includes building a brand image for shipping services, providing B2C services, utilizing the national image for the shipping industry, and participating in international cooperation projects. The text contact approach includes B2B advertising, advertisements focused on the national economic and multifaceted role of the shipping industry, package advertisements for the shipping industry and related industries, the Internet and high-technology media, government-initiated PR activities regarding the multifaceted role of the shipping industry, and funding for advertising the shipping industry.

Scene Text Recognition Performance Improvement through an Add-on of an OCR based Classifier (OCR 엔진 기반 분류기 애드온 결합을 통한 이미지 내부 텍스트 인식 성능 향상)

  • Chae, Ho-Yeol;Seok, Ho-Sik
    • Journal of IKEEE
    • /
    • v.24 no.4
    • /
    • pp.1086-1092
    • /
    • 2020
  • An autonomous agent for real world should be able to recognize text in scenes. With the advancement of deep learning, various DNN models have been utilized for transformation, feature extraction, and predictions. However, the existing state-of-the art STR (Scene Text Recognition) engines do not achieve the performance required for real world applications. In this paper, we introduce a performance-improvement method through an add-on composed of an OCR (Optical Character Recognition) engine and a classifier for STR engines. On instances from IC13 and IC15 datasets which a STR engine failed to recognize, our method recognizes 10.92% of unrecognized characters.

Abstruseness of Rimbaud's Barbare : Autotextuality and Meaning (랭보의 「야만」의 난해성 : '자기텍스트성'과 '의미')

  • Shin, Ok-Keun
    • Cross-Cultural Studies
    • /
    • v.43
    • /
    • pp.327-354
    • /
    • 2016
  • Rimbaud's prose poem, Barbare in Illuminations, is known for its abstruseness with regard to forms, themes, metaphors. This paper first analyzes the poem's grammatical structure to make sense of such an inscrutable piece of work, then discusses its autotextuality in order to decipher its meaning by comparison with Rimbaud's other works. Autotextuality, a method of literary interpretation of Rimbaud's prose poem presented by Steve Murphy, refers to the intertextuality between the author's works. Despite some previous researches focusing on the intertextuality of Barbare, previous authors have failed not only to find its meaning but also to determine its significance. The abstruseness of Rimbaud's Barbare is sometimes considered an example of the meaningless of Rimbaud's work. However, examining the textual structure and the autotextuality builds meaning, rather than rendering the work meaningless. Barbare which consists entirely of noun phrases and metaphors means destruction, fusion and the pure power of regeneration in the original context of Rimbaud's work. This poem is Rimbaud's answer to Baudelaire's poetic question, Any of where out of World, and presents a strange scenery that uses 'the eternal female voice' to reach the Vulcan in the North Pole. Interpretation of Barbare could provide a methodology for reading the difficult Illuminations. The kind of analyses used are, for example, analysis of the text, analysis of verbal indicators, autotextuality, and an understanding of the joy and the solitude in the silence of the poem. Understanding Barbare may provide a method of interpreting the abstruseness of Illuminations. Through this approach, we can connect and combine every fragment of the Illuminations, so that we can reconstruct the story and the adventure contained therein.

Detection of Hidden Knowledge Using a Citation-Based Approach Based on Swanson's ABC Model (인용 정보를 고려한 미발견 공공 지식 추출: Swanson의 ABC 모델 재현 및 확장)

  • Hahm, Jung Eun;Song, Min
    • Journal of the Korean Society for information Management
    • /
    • v.32 no.2
    • /
    • pp.87-103
    • /
    • 2015
  • It is useful to find something valuable for researching through literature based discovery. Swanson's ABC model, known as literature based discovery, suggests the relationship between entities undiscovered yet. This study tries to find the valid relationship between entities by referring to citation which connects articles on similar topic. We collect citation from references in articles, and extract important concepts in titles and abstracts through text mining techniques. We reproduce the relationship between fish oil and Raynaud's disease, which is known as one of Swanson's works, and compare the results with entities identified from traditional approach.

Ensemble-based Counterfeit Detection Algorithm (앙상블 기반의 위조 탐지 알고리즘)

  • Ilkin Taghiyev;Youngbok-Cho
    • Proceedings of the Korean Society of Computer Information Conference
    • /
    • 2023.01a
    • /
    • pp.101-102
    • /
    • 2023
  • 본 연구에서는 인터넷 상에서 발생되는 부정행위를 탐지할수 있는 신뢰 모델을 생성하고 개인의 프라이버시를 보장할수 있는 모델을 제시하였다. 인터넷 상에 게시판에 올려진 부정해위를 탐지하기 위해 앙상블 접근 방식 기반의 분류 모델을 제시하고 자동화된 도구를 제안하였다. 본 연구는 데이터에 대한 탐색적 데이터 분석을 수행하고 얻은 통찰력을 사용해 자연어처리 가반 텍스트를 기반으로 앙상블 기반의 위조 탐지 알고리즘을 제안하였다. 제안 알고리즘의 정확도는 99%로 자연어 처리에 높은 탐지율을 보였다.

  • PDF

A Global-Interdependence Pairwise Approach to Entity Linking Using RDF Knowledge Graph (개체 링킹을 위한 RDF 지식그래프 기반의 포괄적 상호의존성 짝 연결 접근법)

  • Shim, Yongsun;Yang, Sungkwon;Kim, Hong-Gee
    • KIPS Transactions on Software and Data Engineering
    • /
    • v.8 no.3
    • /
    • pp.129-136
    • /
    • 2019
  • There are a variety of entities in natural language such as people, organizations, places, and products. These entities can have many various meanings. The ambiguity of entity is a very challenging task in the field of natural language processing. Entity Linking(EL) is the task of linking the entity in the text to the appropriate entity in the knowledge base. Pairwise based approach, which is a representative method for solving the EL, is a method of solving the EL by using the association between two entities in a sentence. This method considers only the interdependence between entities appearing in the same sentence, and thus has a limitation of global interdependence. In this paper, we developed an Entity2vec model that uses Word2vec based on knowledge base of RDF type in order to solve the EL. And we applied the algorithms using the generated model and ranked each entity. In this paper, to overcome the limitations of a pairwise approach, we devised a pairwise approach based on comprehensive interdependency and compared it.